Persian Scientific Question Answering: A Transformer-based Approach with a Domain-Specific Dataset

سال انتشار: 1404
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 5

فایل این مقاله در 7 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

CSCG06_205

تاریخ نمایه سازی: 4 مهر 1405

چکیده مقاله:

The rapid increase in Persian scientific publications has created a growing need for intelligent systems capable of retrieving precise information from academic texts. This study presents a transformer-based Question Answering (QA) model specifically designed for Persian scientific articles. The proposed system constructs a SQUAD-style extractive QA dataset and fine-tunes the HooshvareLab bert-fa-base-uncased model for Persian QA tasks. The model is trained on an existing general-purpose Persian QA dataset and evaluated on a custom domain-specific test set of ۱۰۰ QA pairs derived from peer-reviewed Persian scientific abstracts. The workflow includes dataset preprocessing, tokenization, model fine-tuning, and evaluation using standard metrics such as EM and F۱ score. Experimental results demonstrate that the fine-tuned model achieves a validation F۱ of ۴۸.۴۵% and EM of ۲۹.۳۰%, while attaining an F۱ of ۴۰.۷۵% and EM of ۲۰.۰۰% on unseen domain-specific test data. These results highlight the model's ability to capture contextual and semantic relationships in Persian scientific texts, despite challenges such as limited annotated data and morphological complexity. The findings suggest that domain-specific evaluation, combined with appropriate fine-tuning, can significantly enhance the performance of Persian QA systems. This work provides a foundational step toward developing advanced Persian-language QA systems for academic and research applications.

نویسندگان

Hediyeh Mosafer

dept. Computer Engineering, University of Guilan

Seyyed Abdorreza Hesam Mohseni

dept. Computer Engineering, University of Guilan, University Lecturer