CIVILICA We Respect the Science
(ناشر تخصصی کنفرانسهای کشور / شماره مجوز انتشارات از وزارت فرهنگ و ارشاد اسلامی: ۸۹۷۱)

A Signal Processing Method for Text Language Identification

عنوان مقاله: A Signal Processing Method for Text Language Identification
شناسه ملی مقاله: JR_IJE-34-6_004
منتشر شده در در سال 1400
مشخصات نویسندگان مقاله:

H. Hassanpour - Image Processing & Data Mining Lab, Shahrood University of Technology, Shahrood, Iran
M. M. AlyanNezhadi - Department of Mathematics, University of Science and Technology of Mazandaran, Behshahr, Iran
M. Mohammadi - Department of Information Technology, Lebanese Frebch, University, Erbil, Kurdistan Region, Iraq

خلاصه مقاله:
Language identification is a critical step prior to any natural language processing. In this paper, a signal processing method for Language Identification is proposed. Sequence of characters in a word and the order of words in stream identify the language. The sequence of characters in a stream provides a signature to recognize the language without understanding its meaning. The signature can be extracted using signal processing techniques via converting texts into time series. Although several research and commercial software have been developed to identify text language, they need a standard dictionary for each language. We proposed a dictionary independent method consisting of three main steps, I) preprocessing, II) clustering and finally III) classification. First, the texts are converted to time series using UTF-۸ codes. Second, to group similar languages, the obtained series are clustered. Third, each cluster is decomposed into ۳۲ sub-bands using a Wavelet packet, and ۳۲ features are extracted from each sub-band. Also, a multilayer perceptron neural network is used to classify the extracted features. The proposed method was tested on our dataset with ۳۱۰۰۰ texts from ۳۱ different languages. The proposed method achieved ۷۲.۲۰% accuracy for language identification.

کلمات کلیدی:
Language Identification, Signal processing, Wavelet Packet Transform, Artificial Neural Network

صفحه اختصاصی مقاله و دریافت فایل کامل: https://civilica.com/doc/1224245/