Publications

(2026). PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors. In Proc. Interspeech 2026.
(2026). Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations. In Proc. Interspeech 2026.
(2025). BitTTS: 1.58-bit 量子化と重みインデキシングによる軽量なテキスト音声合成. 日本音響学会 2025年秋季研究発表会.
(2025). SLASH: 信号処理と自己教師あり学習を組み合わせた基本周波数推定法. 日本音響学会 2025年秋季研究発表会.
(2025). SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch. In Proc. Interspeech 2025.
(2025). Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments. In Proc. Interspeech 2025.
(2025). BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing. In Proc. Interspeech 2025.
(2025). Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control. In Proc. ICASSP2025.
(2024). LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning. In Proc. Interspeech 2024.
(2023). PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions. In Proc. ICASSP2024.