بهره مندی از منابع داده ای چند زبانی برای یادگیری انتقالی در تشخیص گفتمان ناپسند فارسی

فایل این در 95 صفحه با فرمت PDF قابل دریافت می باشد

  • من نویسنده این مقاله هستم

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این :

چکیده :

در این پایان نامه، مسئله تشخیص گفتمان ناپسند در شبکه های اجتماعی مورد بررسی قرار می گیرد. باتوجه به اهمیت ارتباطات آنلاین و نیاز به حفظ یک محیط سالم و متناسب با استانداردهای اخلاقی، هدف اصلی این تحقیق توسعه یک روش دقیق و کارآمد برای تشخیص پیام های گفتمانی ناپسند در شبکه های اجتماعی است. باتوجه به اهمیت و افزایش استفاده از گفتمان ناپسند، در پی رواج استفاده از شبکه های اجتماعی، تحقیقات متعددی برای شناسایی انواع گفتمان ناپسند در انواع زبان ها صورت گرفته است. در مورد زبان فارسی، باتوجه به اینکه به عنوان یک زبان کم منبع شناخته می شود، تحقیقات معدودی انجام شده است. در این پایان نامه بر آنیم تا به مسئله تشخیص اتوماتیک گفتمان ناپسند در زبان فارسی بپردازیم. دراین خصوص به خزش متون شبکه اجتماعی و برچسب گذاری و آماده سازی یک مجموعه دادگان پرداختیم. همچنین از برخی از دادگان موجود در زبان فارسی استفاده نمودیم. در زمینه مدل، باتوجه به کم منبع بودن زبان فارسی، به سراغ مدل های مبتنی بر یادگیری انتقالی (به طور خاص انتقال بین زبانی) و استفاده از منابع داده ای چندزبانه خواهیم پرداخت. با بهره گیری از یک مجموعه دادگان بزرگ گفتمان ناپسند در زبان انگلیسی، به معرفی سه معماری مبتنی بر یادگیری انتقالی خواهیم پرداخت. دو معماری، از روش های ترجمه بین زبانی بهره خواهند برد. معماری سوم نیز از مدل های زبانی چندزبانه مانند M-BERT و XLM استفاده خواهد نمود. نتایج به دست آمده از تست معماری های پیشنهادی حاکی از آن است که این روش ها نسبت به روش های دیگر مبتنی بر یادگیری ماشین و یادگیری عمیق (مانند شبکه های عصبی بازگشتی، شبکه های عصبی پیچشی و استفاده از مدل های زبانی تک زبانه مانند BERT فارسی) بادقت بالاتری (حدود 5 تا 10 درصد) به تشخیص گفتمان ناپسند می پردازد.

نویسندگان

مرتضی نوروزی

هنر آموز کامپیوتر و معاون فنی

فاطمه چعفری نژاد

استاد راهنما

مراجع و منابع این :

لیست زیر مراجع و منابع استفاده شده در این را نمایش می دهد. این مراجع به صورت کاملا ماشینی و بر اساس هوش مصنوعی استخراج شده اند و لذا ممکن است دارای اشکالاتی باشند که به مرور زمان دقت استخراج این محتوا افزایش می یابد. مراجعی که مقالات مربوط به آنها در سیویلیکا نمایه شده و پیدا شده اند، به خود لینک شده اند :
  • [1] Z. Wang, J. Liao, Q. Cao, H. Qi, and Z. ...
  • [2] I. Goodfellow, Y. Bengio, and A. Courville, Deep learning. MIT ...
  • [3] H.-Y. Liao, K.-Y. Chen, and D.-R. Liu, “Virtual Friend Recommendations ...
  • [4] J. Uyheng, D. Bellutta, and K. M. Carley, “Bots Amplify ...
  • [5] T. M. Mitchell, “Machine learning.” 1997. ...
  • [6] S. J. Russell, Artificial intelligence a modern approach. Pearson Education, ...
  • [7] I. Goodfellow, Y. Bengio, and A. Courville, Deep learning. MIT ...
  • [8] M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A ...
  • [9] T. Zhang, “Solving Large Scale Linear Prediction Problems Using Stochastic ...
  • [10] V. Nair and G. E. Hinton, “Rectified linear units improve ...
  • [11] H. Schütze, C. D. Manning, and P. Raghavan, Introduction to ...
  • [12] P.-S. Huang, X. He, J. Gao, L. Deng, A. Acero, ...
  • [13] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput, ...
  • [14] N. Djuric, J. Zhou, R. Morris, M. Grbovic, V. Radosavljevic, ...
  • [15] “6”. ...
  • [16] Z. Waseem and D. Hovy, “Hateful symbols or hateful people? ...
  • [17] B. Gambäck and U. K. Sikdar, “Using Convolutional Neural Networks ...
  • [18] K. Henderson et al., “Metric forensics: a multi-level approach for ...
  • [19] A. M. Founta et al., “Large Scale Crowdsourcing and Characterization ...
  • [20] A. Vaswani et al., “Attention is All you Need,” in ...
  • [21] T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” in ...
  • [22] A. Vaswani et al., “Attention is all you need,” Adv ...
  • [23] K. Cho et al., “Learning phrase representations using RNN encoder-decoder ...
  • [24] X. Zhou, E. Yilmaz, Y. Long, Y. Li, and H. ...
  • [25] S. Feng, S. Liu, M. Li, and M. Zhou, “Implicit ...
  • [26] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: ...
  • [27] F. Carvalho and C. Castro, “Empirical Analysis on the State ...
  • [28] “scholar”. ...
  • [29] P.-S. Huang, X. He, J. Gao, L. Deng, A. Acero, ...
  • [30] “scholar (1)) . ...
  • [31] J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods ...
  • [32] G. Di Santo, R. McCreadie, C. Macdonald, and I. Ounis, ...
  • [33] S. Ling et al., “Adapting Large Language Model with Speech ...
  • [34] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: ...
  • [35] R. Dietrich, M. Opper, and H. Sompolinsky, “Statistical mechanics of ...
  • [36] H. Yenala, M. Chinnakotla, and J. Goyal, “Convolutional bi-directional LSTM ...
  • [37] B. Vandersmissen, “Automated detection of offensive language behavior on social ...
  • [38] Z. Xu and S. Zhu, “Filtering offensive language in online ...
  • [39] M. Bojkovský and M. Pikuliak, “STUFIIT at SemEval-2019 Task 5: ...
  • [40] H. Yenala, M. Chinnakotla, and J. Goyal, “Convolutional bi-directional LSTM ...
  • [41] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” the ...
  • [42] G. Lample and A. Conneau, “Cross-lingual Language Model Pretraining,” CoRR, ...
  • [43] G. Xiang, B. Fan, L. Wang, J. Hong, and C. ...
  • [44] L. Tang and H. Liu, Community detection and mining in ...
  • [45] P. J. Werbos, “Backpropagation through time: what it does and ...
  • [46] A. Chuklin and A. Lavrentyeva, “Adult query classification for web ...
  • [47] S. Whiting and J. M. Jose, “Recent and robust query ...
  • [48] M. Tasgin, A. Herdagdelen, and H. Bingol, “Community detection in ...
  • [49] T. Davidson, D. Warmsley, M. Macy, and I. Weber, “Automated ...
  • [50] W. A. Rodrigues Jr, “From the Photon to Maxwell Equation. ...
  • [51] C. Nobata, J. Tetreault, A. Thomas, Y. Mehdad, and Y. ...
  • [52] A. Cartas, C. Ballester, and G. Haro, “A Graph-Based Method ...
  • [53] Z. Waseem and D. Hovy, “Hateful symbols or hateful people? ...
  • [54] S. Zannettou et al., “The web centipede: understanding how web ...
  • [55] S. Zannettou et al., “The web centipede: understanding how web ...
  • [56] A. See, P. J. Liu, and C. D. Manning, “Get ...
  • [57] Y. Liu et al., “Roberta: A robustly optimized bert pretraining ...
  • [58] N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong, ...
  • [59] A. Wasala, J. Buckley, R. Schäler, and C. Exton, “An ...
  • [60] “10.1162/jmlr.2003.3.4-5.993,” CrossRef Listing of Deleted DOIs, vol. 1, 2000, doi: ...
  • [61] M. Azaouzi, D. Rhouma, and L. Romdhane, “Community detection in ...
  • [62] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” the ...
  • [63] S. MacAvaney, H. R. Yao, E. Yang, K. Russell, N. ...
  • [64] A. Vaswani et al., “Attention is all you need,” Adv ...
  • [65] H. Zhang et al., “Improving Speech Translation by Cross-Modal Multi-Grained ...
  • [66] X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, ...
  • [67] S. Ruder, “A survey of cross-lingual embedding models,” CoRR, vol. ...
  • [68] A. Conneau, G. Lample, M. Ranzato, L. Denoyer, and H. ...
  • [69] M. Artetxe, G. Labaka, E. Agirre, and K. Cho, “Unsupervised ...
  • [70] C. K. Quah, Translation and technology. Springer, 2006. ...
  • [71] A. Wasala, J. Buckley, R. Schäler, and C. Exton, “An ...
  • [72] “mBERT.” ...
  • [73] “XLM.” ...
  • [74] A. Conneau et al., “Unsupervised Cross-lingual Representation Learning at Scale,” ...
  • [75] F. Alkomah and X. Ma, “A Literature Review of Textual ...
  • PerBOLD: A Big Dataset of Persian Offensive language on Instagram Comments [مقاله ژورنالی]
  • نمایش کامل مراجع