跳至主導覽 跳至搜尋 跳過主要內容

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

研究成果: 書貢獻/報告類型會議論文篇章

摘要

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy. Misalignment between these two modalities often reduces the reliability of multimodal emotion recognition systems, particularly in conversational contexts. To address this issue, we introduce an alignment pipeline utilizing pre-trained ASR and speaker diarization models, systematically synchronizing timestamps to generate accurately labeled speaker segments. Our multimodal approach combines textual embeddings extracted via RoBERTa with audio embeddings from Wav2Vec, leveraging cross-attention fusion enhanced by a gating mechanism. Experimental evaluations on the IEMOCAP benchmark dataset demonstrate that precise timestamp alignment improves SER accuracy, outperforming baseline methods that lack synchronization. The results highlight the critical importance of temporal alignment, demonstrating its effectiveness in enhancing overall emotion recognition accuracy and providing a foundation for robust multimodal emotion analysis.

原文英語
主出版物標題Proceedings of 2025 International Conference on Asian Language Processing, IALP 2025
編輯Lei Wang, Rong Tong, Sarah Flora Samson Juan, Yanfeng Lu, Ping Ping Tan, Suhaila Saee, Minghui Dong
發行者Institute of Electrical and Electronics Engineers Inc.
頁面85-90
頁數6
ISBN(電子)9798331589790
DOIs
出版狀態已發佈 - 2025
事件29th International Conference on Asian Language Processing, IALP 2025 - Sarawak, 马来西亚
持續時間: 2025 8月 42025 8月 6

出版系列

名字Proceedings of 2025 International Conference on Asian Language Processing, IALP 2025

會議

會議29th International Conference on Asian Language Processing, IALP 2025
國家/地區马来西亚
城市Sarawak
期間2025/08/042025/08/06

ASJC Scopus subject areas

  • 人工智慧
  • 電腦科學應用
  • 電腦視覺和模式識別
  • 訊號處理

指紋

深入研究「Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization」主題。共同形成了獨特的指紋。

引用此