TY - JOUR
T1 - Flexible VAD-PVAD Transition
T2 - 26th Interspeech Conference 2025
AU - Yu, En Lun
AU - Wang, Chien Chun
AU - Hung, Jeih Weih
AU - Huang, Shih Chieh
AU - Chen, Berlin
N1 - Publisher Copyright:
© 2025 International Speech Communication Association. All rights reserved.
PY - 2025
Y1 - 2025
N2 - In this paper, we propose Flexible Dynamic Encoder RNN (FDE-RNN), an innovative model capable of seamlessly switching between VAD and PVAD without incurring redundant resource consumption. In static PVAD modeling, performing VAD typically requires either merging categories or omitting speaker embeddings, often resulting in excessively large models that are impractical for VAD tasks. In contrast, FDE-RNN efficiently adapts by removing the personalization module when functioning as VAD, significantly reducing resource demands. Furthermore, on PVAD tasks, FDE-RNN leverages dynamic neural networks with a gating-based skipping mechanism, enabling it to bypass redundant computations during non-speech segments, further optimizing computational efficiency. Extensive experiments demonstrate that FDE-RNN outperforms all other prior arts on both PVAD and VAD tasks in terms of overall performance. Notably, when functioning as a VAD, FDE-RNN merely utilizes 30% of the parameters required by the competitive models, underscoring its remarkable efficiency and scalability.
AB - In this paper, we propose Flexible Dynamic Encoder RNN (FDE-RNN), an innovative model capable of seamlessly switching between VAD and PVAD without incurring redundant resource consumption. In static PVAD modeling, performing VAD typically requires either merging categories or omitting speaker embeddings, often resulting in excessively large models that are impractical for VAD tasks. In contrast, FDE-RNN efficiently adapts by removing the personalization module when functioning as VAD, significantly reducing resource demands. Furthermore, on PVAD tasks, FDE-RNN leverages dynamic neural networks with a gating-based skipping mechanism, enabling it to bypass redundant computations during non-speech segments, further optimizing computational efficiency. Extensive experiments demonstrate that FDE-RNN outperforms all other prior arts on both PVAD and VAD tasks in terms of overall performance. Notably, when functioning as a VAD, FDE-RNN merely utilizes 30% of the parameters required by the competitive models, underscoring its remarkable efficiency and scalability.
KW - Dynamic Neural Networks
KW - Personalized Voice Activity Detection
KW - Voice Activity Detection
UR - https://www.scopus.com/pages/publications/105020038629
UR - https://www.scopus.com/pages/publications/105020038629#tab=citedBy
U2 - 10.21437/Interspeech.2025-322
DO - 10.21437/Interspeech.2025-322
M3 - Conference article
AN - SCOPUS:105020038629
SN - 2308-457X
SP - 5793
EP - 5797
JO - Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
JF - Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Y2 - 17 August 2025 through 21 August 2025
ER -