a Robust Sound Source Localization Using CNN-LSTM and Q-Learning in Dynamic Scenes

سال انتشار: 1404
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 28

فایل این مقاله در 12 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

ISAV15_062

تاریخ نمایه سازی: 7 مرداد 1405

چکیده مقاله:

This paper proposes a novel hybrid model for localizing moving sound sources using microphone arrays, integrating convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and Q-Learning. The proposed method incorporates multidimensional acoustic features namely, Mel-frequency cepstral coefficients (MFCCs) and Mel-spectrograms-extracted from time-frequency audio signals into a Q-Learning framework that leverages CNN and LSTM architectures. This integration effectively mitigates the challenges posed by noise and reverberation in complex acoustic environments. For performance evaluation, the model utilizes synthetic data generated by the Room Impulse Response Generator software. Experimental designs were carefully constructed to assess the robustness of the model, with particular emphasis on the impact of incrementally increasing noise levels and reverberation times on localization performance. The results demonstrate that the proposed algorithm significantly outperforms existing baseline methods, achieving a localization accuracy of ۹۶% and a root mean square error (RMSE) of ۳.۶۵۰° in direction-of-arrival (DOA) estimation. These findings underscore the substantial potential of the model to deliver reliable and efficient sound source localization in real-world acoustic scenarios.

کلیدواژه ها:

نویسندگان

Elham Yazdankhah

Ph.D. Student, Department of Electrical Engineering, Lorestan University, Khorramabad, Lorestan, Iran.

Salman Karimi

Assistant Professor, Department of Electrical Engineering, Lorestan University, Khorramabad, Lorestan, Iran.