a Robust Sound Source Localization Using CNN-LSTM and Q-Learning in Dynamic Scenes
محل انتشار: پانزدهمین کنفرانس بین المللی آکوستیک و ارتعاشات
سال انتشار: 1404
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 28
فایل این مقاله در 12 صفحه با فرمت PDF قابل دریافت می باشد
- صدور گواهی نمایه سازی
- من نویسنده این مقاله هستم
استخراج به نرم افزارهای پژوهشی:
شناسه ملی سند علمی:
ISAV15_062
تاریخ نمایه سازی: 7 مرداد 1405
چکیده مقاله:
This paper proposes a novel hybrid model for localizing moving sound sources using microphone arrays, integrating convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and Q-Learning. The proposed method incorporates multidimensional acoustic features namely, Mel-frequency cepstral coefficients (MFCCs) and Mel-spectrograms-extracted from time-frequency audio signals into a Q-Learning framework that leverages CNN and LSTM architectures. This integration effectively mitigates the challenges posed by noise and reverberation in complex acoustic environments. For performance evaluation, the model utilizes synthetic data generated by the Room Impulse Response Generator software. Experimental designs were carefully constructed to assess the robustness of the model, with particular emphasis on the impact of incrementally increasing noise levels and reverberation times on localization performance. The results demonstrate that the proposed algorithm significantly outperforms existing baseline methods, achieving a localization accuracy of ۹۶% and a root mean square error (RMSE) of ۳.۶۵۰° in direction-of-arrival (DOA) estimation. These findings underscore the substantial potential of the model to deliver reliable and efficient sound source localization in real-world acoustic scenarios.
کلیدواژه ها:
نویسندگان
Elham Yazdankhah
Ph.D. Student, Department of Electrical Engineering, Lorestan University, Khorramabad, Lorestan, Iran.
Salman Karimi
Assistant Professor, Department of Electrical Engineering, Lorestan University, Khorramabad, Lorestan, Iran.