Feature Selection in Big Data Using Parallel Processing of Map-Reduce Model in Hadoop Platform Based on Clustering and Optimization Algorithms

سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 29

فایل این مقاله در 12 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

UTCONF10_047

تاریخ نمایه سازی: 26 شهریور 1405

چکیده مقاله:

Data platforms that are large in size, despite the opportunities they create, pose many computational challenges. One of the problems with large data sets is that most of the time all the features of the data are not critical to finding the knowledge that lies in the data. For this reason, in many areas, data reduction is one of the most significant issues. One efficient way to reduce data dimensions is to use feature selection. During the attribute selection process, a subset of the primary attributes is selected by removing the irrelevant and redundant attributes. In this paper, for the first time, a new hierarchical algorithm based on feature relationships for feature selection in the Hadoop context is presented. In this paper, for the first time, the ability to explore and search in the standard optimization algorithm will be strengthened by presenting a new and improved version of the particle optimization algorithm that will have intersection and mutation operators. Then, using this new optimization algorithm and utilizing feature clustering, and in selecting the final features using the node centrality criterion, a new method of feature selection in the Hadoop context is presented. The feature clustering method presented in this paper will be completely different from the previous algorithms and instead of using traditional clustering models, it will form the final clusters based on the graphical structure of the features and the relationships between the features. In the new model of particle optimization algorithms, which use new operators to improve the search capability, when the parameter ۰ is equal to ۰.۴ or ۰.۶ or the parameter w is equal to ۲ or ۳, the algorithm has higher classification accuracy.

نویسندگان

Sara Dehghani

Ph.D. candidate, Department of Computer Engineering, Yas.C., Islamic Azad University, Yasuj, Iran.

Razieh Malekhosseini

Ph.D. Assistant professor, Department of Computer Engineering, Yas.C., Islamic Azad University, Yasuj, Iran.

Karamollah Bagherifard

Ph.D. Associate professor, Department of Computer Engineering, Yas.C., Islamic Azad University, Yasuj, Iran.

S. Hadi Yaghoubyan

Ph.D. Assistant professor, Department of Computer Engineering, Yas.C., Islamic Azad University, Yasuj, Iran.