跳至主導覽 跳至搜尋 跳過主要內容

Flexible VAD-PVAD Transition: A Detachable PVAD Module for Dynamic Encoder RNN VAD

  • En Lun Yu
  • , Chien Chun Wang
  • , Jeih Weih Hung
  • , Shih Chieh Huang
  • , Berlin Chen

研究成果: 雜誌貢獻會議論文同行評審

2   連結會在新分頁中打開 引文 斯高帕斯(Scopus)

摘要

In this paper, we propose Flexible Dynamic Encoder RNN (FDE-RNN), an innovative model capable of seamlessly switching between VAD and PVAD without incurring redundant resource consumption. In static PVAD modeling, performing VAD typically requires either merging categories or omitting speaker embeddings, often resulting in excessively large models that are impractical for VAD tasks. In contrast, FDE-RNN efficiently adapts by removing the personalization module when functioning as VAD, significantly reducing resource demands. Furthermore, on PVAD tasks, FDE-RNN leverages dynamic neural networks with a gating-based skipping mechanism, enabling it to bypass redundant computations during non-speech segments, further optimizing computational efficiency. Extensive experiments demonstrate that FDE-RNN outperforms all other prior arts on both PVAD and VAD tasks in terms of overall performance. Notably, when functioning as a VAD, FDE-RNN merely utilizes 30% of the parameters required by the competitive models, underscoring its remarkable efficiency and scalability.

原文英語
頁(從 - 到)5793-5797
頁數5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
DOIs
出版狀態已發佈 - 2025
事件26th Interspeech Conference 2025 - Rotterdam, 荷兰
持續時間: 2025 8月 172025 8月 21

ASJC Scopus subject areas

  • 軟體
  • 訊號處理
  • 語言與語言學
  • 建模與模擬
  • 人機介面

指紋

深入研究「Flexible VAD-PVAD Transition: A Detachable PVAD Module for Dynamic Encoder RNN VAD」主題。共同形成了獨特的指紋。

引用此