A Survey of Part of Speech Tagging of Latin and non-Latin Script Languages: A More Vivid View on Persian

سال انتشار: 1400
نوع سند: مقاله ژورنالی
زبان: انگلیسی
مشاهده: 351

فایل این مقاله در 16 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

JR_LART-6-1_004

تاریخ نمایه سازی: 9 آبان 1401

چکیده مقاله:

Thir research is a general overview of the Latin Script languages part of speech (POS) tagging with a specific focus on the non-Latin script languages, especially Persian. The study reviews the progress in POS tagging among the ۲۳ highest native spoken languages in the world. Some of these languages follow the right-to-left (RTL) writing system such as Arabic, Urdu and Persian which have their own specific issues in POS tagging. This paper also goes through the issues and challenges which occurs during the tokenization and part of speech tagging of these languages. The challenges can be common between the languages or be specified to one. The Persian Language is chosen as the main interest of this paper and an attempt is made to critically overview the recent studies on Persian part of speech tagging and enumerate the specific challenges occurring in these studies. Reviewing the bulk of literature and examining the features, that significant challenges of the researches on Persian were generally in the tokenization level and mostly as a result of using the Arabic script and its characteristics.

کلیدواژه ها:

نویسندگان

Mesian Moghaddam

Assistant professor, Facullty of science, Fasa University, Iran

Niloofar Jafarpour

M.A.Institute of Linguistics and Literay Studies Technical University of Darmstadt, Germany