Kernelled Naïve Bayes Using a Balanced Dataset for Accurate Classification of the Material Toxicity

سال انتشار: 1400
نوع سند: مقاله ژورنالی
زبان: انگلیسی
مشاهده: 298

فایل این مقاله در 14 صفحه با فرمت PDF قابل دریافت می باشد

این مقاله در بخشهای موضوعی زیر دسته بندی شده است:

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

JR_AJCS-4-2_006

تاریخ نمایه سازی: 31 فروردین 1400

چکیده مقاله:

In this work, a new multi-class classification approach was employed in the QSAR model to assess chemical toxicity prediction through handling the imbalanced dataset as the critical preprocessing step in the training dataset. Various classifiers of the decision tree, K-NN, naïve Bayes, kernelled naïve Bayes, and SVM and two distinct acute aquatic toxicity datasets towards Daphnia Magna and Fathead Minnow Fish were used to evaluate the generality of the approach. The quantitative response (LC50) was discretized into ten bins. Imbalanced dataset classification leads to a high level of errors since the classifier tends to learn from the majority class more than the minority class. Each training dataset was specified by different weights related to the class population. These datasets were then bootstrapped based on their weights to convert the imbalanced dataset into a balanced one. This approach enhanced the accuracy of classification of material toxicity dramatically (up to 99%). Balanced dataset classification had high overall accuracy when correlated attributes were removed. Therefore, fewer attributes are sufficient to predict material toxicity.The overall accuracy improvement of the decision tree, K-NN, naïve Bayes, kernelled naïve Bayes, and SVM for the Daphnia Magna dataset after balancing the data set are 58.03%, 55.08%, 9.09%, 72.48%, and 53.05%, respectively.

نویسندگان

Ali Ekramipooya

Department of Chemical and Petroleum Engineering, Sharif University of Technology, Tehran, Iran

Davood Rashtchian

Department of Chemical and Petroleum Engineering, Sharif University of Technology, Tehran, Iran

Mehrdad Boroushaki

Department of Energy Engineering, Sharif University of Technology, Tehran, Iran. P.O.Box ۱۴۵۶۵-۱۱۴

مراجع و منابع این مقاله:

لیست زیر مراجع و منابع استفاده شده در این مقاله را نمایش می دهد. این مراجع به صورت کاملا ماشینی و بر اساس هوش مصنوعی استخراج شده اند و لذا ممکن است دارای اشکالاتی باشند که به مرور زمان دقت استخراج این محتوا افزایش می یابد. مراجعی که مقالات مربوط به آنها در سیویلیکا نمایه شده و پیدا شده اند، به خود مقاله لینک شده اند :
  • [1] M. Rausand, Risk assessment: theory, methods, and applications, John ...
  • [2] G. Popov, B.K. Lyon, B. Hollcroft, John Wiley & ...
  • [3] Development (OECD) Staff, Development. Working Party on Environmental Performance, ...
  • [4] CEC, Regulation (EC) No. 1907/2006 of the European Parliament ...
  • [5] M.T. Cronin, J.D. Walker, J.S. Jaworska, M.H. Comber, C.D. ...
  • [6] R. Combes, C. Grindon, M.T. Cronin, D.W. Roberts, J.F. ...
  • [7] M. Cronin, Chemical Toxicity Prediction: Category Formation and Read-Across, ...
  • [8] H. Liu, E. Papa, P. Gramatica, Chem. Res. Toxicol., ...
  • [9] T.M. Mitchell, Machine Learning, McGraw-Hill, 1997. ...
  • [10] D.E. Jones, H. Ghandehari, J.C. Facelli, Comput. Methods Programs ...
  • [11] H. Yang, L. Sun, W. Li, G. Liu, Y. ...
  • [12] A.A. Toropov, A.P, Toropova, I. Raska Jr, D. Leszczynska, ...
  • [13] S. Schmidt, M. Schindler, D. Faber, J. Hager, SAR ...
  • [14] O. Tinkov, V.Y. Grigorev, A.N. Razdolsky, L.D. Grigoryeva, J.C. ...
  • [15] F. Lunghini, G. Marcou, P. Azam, M.H. Enrici, E. ...
  • [16] M. Marzo, G.J. Lavado, F. Como, A.P. Toropova, A.A. ...
  • [17] J. Polanski, Chemoinformatics: From Chemical Art to Chemistry in ...
  • [18] S.R. Kazmi, R. Jun, M.S. Yu, C. Jung, D. Na, ...
  • [19] Y. Wu, G. Wang, Int. J. Mol. Sci., 2018, ...
  • [20] S.J. Russell, P. Norvig, Artificial Intelligence-A Modern Approach, 3rd ...
  • [21] E. Alpaydin, Introduction to machine learning, MIT press, 2020. ...
  • [22] M. Mohri, A. Rostamizadeh, A. Talwalkar, Foundations of Machine ...
  • [23] S. Cassani, S. Kovarich, E. Papa, P.P. Roy, L. ...
  • [24] F. Abbasitabar, V. Zare-Shahabadi, Chemosphere, 2017, 172, 249–259. ...
  • [25] B. Giner, C. Lafuente, D. Lapeña, D. Errazquin, L. ...
  • [26] R. Aalizadeh, C. Peter, N.S. Thomaidis, Environ. Sci. Process ...
  • [27] T. Fan, G. Sun, L. Zhao, X. Cui, R. ...
  • [28] H. Zhang, P. Yu, J.X. Ren, X.B. Li, H.L., ...
  • [29] H. Zhang, C. Shen, R.Z. Liu, J. Mao, C.T. ...
  • [30] A. Lillicrap, S.J. Moe, R. Wolf, K.A. Connors, J.M. ...
  • [31] J. Roy, P.K. Ojha, E. Carnesecchi, A. Lombardo, K. ...
  • [32] N. Abramenko, L. Kustov, L. Metelytsia, V. Kovalishyn, I. ...
  • [33] L. Yang, Y. Wang, J. Chang, Y. Pan, R. ...
  • [34] K. Khan, P.M. Khan, G. Lavado, C. Valsecchi, J. ...
  • [35] S.K. Pandey, P.K. Ojha, K. Roy, Chemosphere, 2020, 252, ...
  • [36] X. Jin, M. Jin, L. Sheng, Comput. Biol. Med., ...
  • [37] J. Han, J. Pei, M. Kamber, Data Mining: Concepts ...
  • [38] M. Kantardzic, Data Mining: Concepts, Models, Methods, and Algorithms, ...
  • [39] A. Pérez, P. Larrañaga, I. Inza. Int. J. Approx. Reason., ...
  • [40] G.H. John, P. Langley, Estimating continuous distributions in Bayesian ...
  • [41] K.P. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press: ...
  • [42] L. Michielan, L. Pireddu, M. Floris, S. Moro, Mol. ...
  • [43] V. Kotu, B. Deshpande, Predictive analytics and data mining: ...
  • [44] N. Japkowicz, S. Stephen, Intell. Data Anal., 2002, 6, 429–449. ...
  • [45] S. Barua, M.M. Islam, X. Yao, K. Murase, IEEETrans. Knowl. Data Eng., ...
  • [46] C.X. Ling, C. Li. Data mining for direct marketing: ...
  • [47] A. Fernández, S. Garcia, F. Herrera, N.V. Chawla, J. ...
  • [48] Z.H. Zhou, X.Y. Liu, Comput. Intell., 2010, 26, 232–257. ...
  • [49] A. Fernández, V. López, M. Galar, M.J. Del Jesus, ...
  • [50] D.M. Hawkins, S.C. Basak, D. Mills, J. Chem. Inform. Comput. ...
  • [51] S.K. Jha, T.H. Yoon, Z. Pan, Comput. Biol. Med., ...
  • [52] A. Rácz, D. Bajusz, K. Héberger, Mol. Inform., 2019, ...
  • [53] R. Todeschini, V. Consonni, Handbook of molecular descriptors, Wiley-VCH: ...
  • [54] M. Cassotti, D. Ballabio, V. Consonni, A. Mauri, I.V. ...
  • [55] M. Cassotti, D. Ballabio, R. Todeschini, V. Consonni, SAR ...
  • [56] D. Dua, C. Graff, UCI Machine Learning Repository. School ...
  • [57] G.E. Jensen, J.R. Niemelä, E.B. Wedebye, N.G. Nikolov, SAR ...
  • [58] M.T. Martin, T.B. Knudsen, D.M. Reif, K.A. Houck, R.S. ...
  • [59] C. Jiang, H. Yang, P. Di, W. Li, Y. ...
  • نمایش کامل مراجع