A Lightweight ECAPA-TDNN Framework for Persian Speech Emotion Recognition

سال انتشار: 1405
نوع سند: مقاله ژورنالی
زبان: انگلیسی
مشاهده: 73

فایل این مقاله در 16 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

JR_IJWR-9-2_004

تاریخ نمایه سازی: 14 مرداد 1405

چکیده مقاله:

Speech emotion recognition (SER) plays an important role in affective computing and human–computer interaction. However, developing robust SER systems for low-resource languages such as Persian remains challenging due to limited annotated data, class imbalance, inter-speaker variability, and acoustic overlap among emotion categories. This study proposes a compact transfer-learning framework for Persian SER based on the ECAPA-TDNN architecture together with a novel Emotion-Aware Loss. By incorporating inter-emotion similarity into the optimization objective, the proposed method improves discrimination between acoustically similar emotion categories without modifying the underlying network architecture. Experiments on the speaker-independent ShEMO benchmark demonstrate that the proposed framework achieves ۸۳.۰۴% Accuracy, ۸۸.۵۴% Unweighted Average Recall (UAR), and ۸۸.۲۴% F۱-score, consistently outperforming ECAPA-TDNN baselines trained with cross-entropy, standard Additive Angular Margin Softmax (AAM-Softmax), and fixed similarity weighting. The results demonstrate that similarity-aware optimization improves emotional representation learning while preserving the compact architecture of ECAPA-TDNN, making the proposed framework an effective approach for Persian speech emotion recognition.

نویسندگان

Mobina Esmaeili

Department of Computer Engineering, Faculty of Engineering, Alzahra University, Tehran, Iran.

Mohammad Raziei

Department of Electrical Engineering, Faculty of Engineering, Sharif University of Technology, Tehran, Iran.

Vajiheh Sabeti

Department of Computer Engineering, Faculty of Engineering, Alzahra University, Tehran, Iran.

مراجع و منابع این مقاله:

لیست زیر مراجع و منابع استفاده شده در این مقاله را نمایش می دهد. این مراجع به صورت کاملا ماشینی و بر اساس هوش مصنوعی استخراج شده اند و لذا ممکن است دارای اشکالاتی باشند که به مرور زمان دقت استخراج این محتوا افزایش می یابد. مراجعی که مقالات مربوط به آنها در سیویلیکا نمایه شده و پیدا شده اند، به خود مقاله لینک شده اند :
  • M. El Ayadi, M. S. Kamel, and F. Karray, “Survey ...
  • M. B. Akçay and K. Oğuz, “Speech emotion recognition: Emotional ...
  • R. A. Khalil, E. Jones, M. I. Babar, T. Jan, ...
  • S. Latif, R. Rana, S. Khalifa, Raja Jurdak, J. Qadir, ...
  • O. M. Nezami, P. J. Lou, and M. Karami, “ShEMO ...
  • D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. ...
  • B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel ...
  • K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive Statistics Pooling ...
  • M. Ravanelli et al., “SpeechBrain: A General-Purpose Speech Toolkit,” arXiv.org, ...
  • A. Keesing, Y. S. Koh, and M. Witbrock, “Acoustic Features ...
  • A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav۲vec ...
  • A. S. Tehrani, N. Faridani, and R. Toosi, “Unsupervised Representations ...
  • A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. ...
  • Emotion Recognition for Persian Speech Using Convolutional Neural Network and Support Vector Machine [مقاله ژورنالی]
  • A. Yazdani, H. Simchi, and Y. Shekofteh, “Emotion Recognition In ...
  • SeyedMilad Ranaei Siadat, I. M. Voronkov, and A. A. Kharlamov, ...
  • M. Shayaninasab and B. Babaali, “Persian Speech Emotion Recognition by ...
  • A. Radford, J. W. Kim, T. Xu, G. Brockman, C. ...
  • A. S. Alluhaidan, O. Saidani, R. Jahangir, M. A. Nauman, ...
  • A. Satt, S. Rozenberg, and R. Hoory, “Efficient Emotion Recognition ...
  • J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: ...
  • M. Esmaeili, S. M. H. Hasheminejad, and V. Sabeti, “A ...
  • D. Ververidis and C. Kotropoulos, “Emotional speech recognition: Resources, features, ...
  • Zhihong Zeng, M. Pantic, G. I. Roisman, and T. S. ...
  • G. Trigeorgis et al., “Adieu features? End-to-end speech emotion recognition ...
  • Panagiotis Tzirakis, J. Zhang, and B. Schuller, “End-to-End Speech Emotion ...
  • S. Mirsamadi, E. Barsoum, and C. Zhang, “Automatic speech emotion ...
  • Y. Zhang, J. Du, Z. Wang, J. Zhang, and Y. ...
  • E. Lieskovská, M. Jakubec, R. Jarina, and M. Chmulík, “A ...
  • E. al. Dhairyashil Patil, “Federated Learning in Real-Time Medical IoT: ...
  • M. Sharma, “Multi-Lingual Multi-Task Speech Emotion Recognition Using wav۲vec ۲.۰,” ...
  • T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A ...
  • T. Gao, X. Yao, and D. Chen, “SimCSE: Simple Contrastive ...
  • L. H. Nguyen, N. T. Pham, M. Khan, A. Othmani, ...
  • R. Liu, Y. Hu, Y. Ren, X. Yin, and H. ...
  • A. Shirian, S. Tripathi, and T. Guha, “Learnable Graph Inception ...
  • A. Shirian and T. Guha, “Compact Graph Architecture for Speech ...
  • Y. Hu, Y. Tang, H. Huang, and L. He, “A ...
  • J. Li, X. Wang, Guorong Lv, and Z. Zeng, “GraphMFT: ...
  • L. Jiang, X. Wang, Guoqing Lv, and Z. Zeng, “GraphCFC: ...
  • B. T. Atmaja, A. Sasou, and M. Akagi, “Survey on ...
  • D. Busbridge, D. Sherburn, P. Cavallo, and N. Y. Hammerla, ...
  • K. Wang, W. Shen, Y. Yang, X. Quan, and R. ...
  • M. Hamidi, “Emotion Recognition from Persian Speech with Neural Network,” ...
  • M. Esmaeili, M. Raziei and V. Sabeti, "ECAPA-TDNN for Persian ...
  • نمایش کامل مراجع