TY - GEN
T1 - Adversarial Learning for Duration Prediction in Indonesian Text-to-Speech
T2 - 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
AU - Wiguna, Yoga Tiara
AU - Prihasto, Bima
AU - Pratama, Boby Mugi
AU - Yeh, Chia Hung
AU - Wang, Jia Ching
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Text-to-Speech (TTS) technology has significantly progressed with deep learning, especially through models like Variational Autoencoder with Adversarial Learning for End-toEnd Text-to-Speech (VITS). However, improving audio quality particularly in duration diversity remains a challenge, especially for languages like Indonesian due to limited datasets and research. This study compares the performance of VITS using Stochastic Duration Predictor (SDP) and Deterministic Duration Predictor (DDP), while also exploring the impact of adversarial training on duration prediction. Evaluation employed subjective Mean Opinion Score (MOS) and objective Cosine Similarity using Resemblyzer. Two datasets were used: 343 formal audio samples and 1250 mixed (formal and informal) samples. The more diverse dataset achieved better results, with a cosine similarity of 0.91124 and a MOS of 4.54. Findings indicate that SDP produces more natural durations, and adversarial learning enhances audio quality through better duration modeling.
AB - Text-to-Speech (TTS) technology has significantly progressed with deep learning, especially through models like Variational Autoencoder with Adversarial Learning for End-toEnd Text-to-Speech (VITS). However, improving audio quality particularly in duration diversity remains a challenge, especially for languages like Indonesian due to limited datasets and research. This study compares the performance of VITS using Stochastic Duration Predictor (SDP) and Deterministic Duration Predictor (DDP), while also exploring the impact of adversarial training on duration prediction. Evaluation employed subjective Mean Opinion Score (MOS) and objective Cosine Similarity using Resemblyzer. Two datasets were used: 343 formal audio samples and 1250 mixed (formal and informal) samples. The more diverse dataset achieved better results, with a cosine similarity of 0.91124 and a MOS of 4.54. Findings indicate that SDP produces more natural durations, and adversarial learning enhances audio quality through better duration modeling.
UR - https://www.scopus.com/pages/publications/105030446567
UR - https://www.scopus.com/pages/publications/105030446567#tab=citedBy
U2 - 10.1109/APSIPAASC65261.2025.11249180
DO - 10.1109/APSIPAASC65261.2025.11249180
M3 - Conference contribution
AN - SCOPUS:105030446567
T3 - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
SP - 1986
EP - 1990
BT - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 22 October 2025 through 24 October 2025
ER -