TY - GEN
T1 - Revealing the Role of Audio Channels in ASR Performance Degradation
AU - Huang, Kuan Tang
AU - Chen, Li Wei
AU - Lee, Hung Shin
AU - Chen, Berlin
AU - Wang, Hsin Min
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences.
AB - Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences.
KW - adapter modules
KW - automatic speech recognition
KW - channel robustness
UR - https://www.scopus.com/pages/publications/105036597884
UR - https://www.scopus.com/pages/publications/105036597884#tab=citedBy
U2 - 10.1109/ASRU65441.2025.11434673
DO - 10.1109/ASRU65441.2025.11434673
M3 - Conference contribution
AN - SCOPUS:105036597884
T3 - ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
BT - ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Y2 - 6 December 2025 through 10 December 2025
ER -