1. SS&SD&ASR TASLP
    Online Frontend System for Multi-Talker DSR Using Neural Blind Source Separation and Diarization
    Yoshiaki Bando, Tomohiko Nakamura, Satoru Fukayama, and Shinji Watanabe
    IEEE/ACM Transactions on Audio, Speech, and Language Processing 2026
  2. ASR OJSP
    Uncertainty-Based Streaming ASR with Evidential Deep Learning
    Hiroaki Sato, Asahi Sakuma, Ryuga Sugano, Tadashi Kumano, Yoshihiko Kawai, Shinji Watanabe, and Tetsuji Ogawa
    IEEE Open Journal of Signal Processing 2026
  3. Multimodal TMM
    TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
    Minsu Kim, Jee-weon Jung, Hyeongseop Rha, Soumi Maiti, Siddhant Arora, Xuankai Chang, Shinji Watanabe, and Yong Man Ro
    IEEE Transactions on Multimedia 2026
  4. SS&ASR CSL
    An End-to-End Integration of Speech Separation and Recognition with Self-Supervised Learning Representation
    Yoshiki Masuyama, Xuankai Chang, Wangyou Zhang, Samuele Cornell, Zhong-Qiu Wang, Nobutaka Ono, Yanmin Qian, and Shinji Watanabe
    Computer Speech & Language 2026
  5. Evaluation EMNLP
    Benchmarking Speech-to-Speech Translation Models
    Alkis Koudounas, Hayato Futami, Quentin Jodelet, Osamu Take, Shinji Watanabe, and Emiru Tsunoo
    In Proceedings of Findings of EMNLP 2026
  6. Speech-LLM EMNLP
    Long Listening Thoughts: Eliciting Open Auditory Reasoning with Deliberative Perception and Cognitive Refinement
    Jaeyeon Kim, Chao-Han Huck Yang, Luoyi Zhang, Chan-Jan Hsu, Fernando Ruiloba Portilla, Jinchuan Tian, Shinji Watanabe, and Carlos Busso
    In Proceedings of Findings of EMNLP 2026
  7. Evaluation EMNLP
    MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
    Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, and David Harwath
    In Proceedings of Findings of EMNLP 2026
  8. ASR&Speech-LLM EMNLP
    Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
    Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, and Carlos Busso
    In Proceedings of Findings of EMNLP 2026
  9. Speech-LLM COLM
    Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
    Jinchuan Tian, Haoran Wang, Bo-Hao Su, Chien-yu Huang, Qingzheng Wang, Jiatong Shi, William Chen, Xun Gong, Siddhant Arora, Chin-Jou Li, Masao Someki, Takashi Maekaku, Yusuke Shinohara, Jin Sakuma, Keita Goto, Chao-Han Huck Yang, and Shinji Watanabe
    In Proceedings of COLM 2026
  10. ASR Interspeech
    YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
    William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  11. ASR&SD Interspeech
    Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
    Alexander Polok, Ivan Medennikov, Jan Černocký, Shinji Watanabe, Lukáš Burget, and Samuele Cornell
    In Proceedings of Interspeech 2026
  12. Speech-LLM&SD Interspeech
    Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning
    Alexander Polok, Samuele Cornell, Sathvik Udupa, Jan Černocký, Shinji Watanabe, and Lukáš Burget
    In Proceedings of Interspeech 2026
  13. Speech-LLM Interspeech
    Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption
    Xun Gong, Jinchuan Tian, Haoran Wang, William Chen, Shinji Watanabe, and Yanmin Qian
    In Proceedings of Interspeech 2026
  14. TTS&Speech-LLM Interspeech
    Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis
    Jinchuan Tian, Haoran Wang, Siddhant Arora, Takashi Maekaku, Keita Goto, Jin Sakuma, Yusuke Shinohara, Chao-Han Huck Yang, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  15. Evaluation Interspeech
    ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling
    Zhuoyan Tao, Jiatong Shi, Hye-jin Shim, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  16. Speech-LLM Interspeech
    An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation
    In Proceedings of Interspeech 2026
  17. ASR Interspeech
    An Empirical Recipe for Universal Phone Recognition
    Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, and David Mortensen
    In Proceedings of Interspeech 2026
  18. ASR&Speaker Interspeech
    Speaker-Aware Hypothesis Clustering and Merging for Target-Speaker-free and Target-Speaker Multi-Talker ASR
    Yosuke Kashiwagi, Osamu Take, Hayato Futami, Emiru Tsunoo, Siddhant Arora, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  19. SE&Evaluation Interspeech
    URGENT-MOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment
    Wei Wang, Wangyou Zhang, Chenda Li, Jiahe Wang, Marvin Sach, Kohei Saijo, Samuele Cornell, Yihui Fu, Zhaoheng Ni, Mengxiao Bi, Tim Fingscheidt, Shinji Watanabe, and Yanmin Qian
    In Proceedings of Interspeech 2026
  20. SSL Interspeech
    Online Predictive Coding for Dual-Mode Self-Supervised Speech Models
    Keita Goto, Takashi Maekaku, Jin Sakuma, Jinchuan Tian, Yusuke Shinohara, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  21. Speech-LLM Interspeech
    Adapting Text LLMs to Speech via Multimodal Depth Up-Scaling
    Kazuki Yano, Jun Suzuki, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  22. Dialogue Interspeech
    Endpoint Anticipation for Low-Latency Spoken Dialogue
    Sathvik Udupa, Shinji Watanabe, Petr Schwarz, and Jan Černocký
    In Proceedings of Interspeech 2026
  23. ASR Interspeech
    ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
    Masao Someki, Alexander Polok, Carlos Carvalho, Chyi-Jiunn Lin, Da-Hee Yang, Jiatong Shi, Jinchuan Tian, Nelson Enrique Yalta Soplin, Samuele Cornell, Siddhant Arora, Francisco Teixeira, Wei Wang, William Chen, Alberto Abad, Chenda Li, Shinji Watanabe, and Wangyou Zhang
    In Proceedings of Interspeech 2026
  24. Evaluation Interspeech
    Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
    Naohiro Tawara, Samuele Cornell, Alexander Polok, Marc Delcroix, Lukáš Burget, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  25. Evaluation Interspeech
    Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings
    Ryo Fukuda, Takatomo Kano, Siddhant Arora, Marc Delcroix, Naohiro Tawara, Atsunori Ogawa, Yuya Chiba, Atsushi Ando, William Chen, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  26. ER&Speech-LLM Interspeech
    Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions
    Abinay Reddy Naini, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, and Carlos Busso
    In Proceedings of Interspeech 2026
  27. SSL&Tokenizer Interspeech
    Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
    Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell, William Chen, Satoru Fukayama, and Shinji Watanabe
    In Proceedings of Interspeech 2026
  28. ASR Interspeech
    Which Data Matter? Embedding-Based Data Selection for Speech Recognition
    Zakaria Aldeneh, Skyler Seto, Maureen Seyssel, Jie Chi, Zijin Gu, Takuya Higuchi, Jee-weon Jung, Shinji Watanabe, David Grangier, Barry-John Theobald, and Tatiana Likhomanenko
    In Proceedings of Interspeech 2026
  29. Speech-LLM ICML
    ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools – From Consensus Learning to Ambiguity-Driven Emotion Reasoning
    Esther Sun, Bo-Hao Su, Abinay Reddy Naini, Shinji Watanabe, and Carlos Busso
    In ICML 2026
  30. Dialogue ICML
    Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
    Siddhant Arora, Haidar Khan, Kai Sun, Xin Luna Dong, Sajal Choudhary, Seungwhan Moon, Xinyuan Zhang, Adithya Sagar, Surya Teja Appini, Kaushik Patnaik, Sanat Sharma, Shinji Watanabe, Anuj Kumar, Ahmed A Aly, Yue Liu, Florian Metze, and Zhaojiang Lin
    In ICML 2026
  31. Evaluation ICML
    LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
    Amir Ivry, and Shinji Watanabe
    In ICML 2026
  32. Speech-LLM ICML
    AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
    William Chen, Prem Seetharaman, Rithesh Kumar, Oriol Nieto, Shinji Watanabe, Justin Salamon, and Zeyu Jin
    In ICML 2026
  33. Evaluation ACL
    Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
    Guan-Ting Lin, Shih-Yun Shan Kuan, Jiatong Shi, Kai-Wei Chang, Siddhant Arora, Shinji Watanabe, and Hung-yi Lee
    In ACL 2026
  34. ASR ACL
    POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
    Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David R. Mortensen, and Shinji Watanabe
    In ACL 2026
  35. Evaluation ACL
    PRiSM: Benchmarking Phone Realization in Speech Models
    Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi, Eunjung Yeo, Ryan Soh-Eun Shim, Hanyu Zhou, Brendon Boldt, Karen Rosero, Kalvin Chang, Darsh Agrawal, Keer Xu, Chao-Han Huck Yang, Jian Zhu, Shinji Watanabe, and David R. Mortensen
    In ACL 2026
  36. SLU ACLFindings
    PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding
    Masao Someki, Chien-yu Huang, Siddhant Arora, Samuele Cornell, Markus Müller, Nathan Susanj, Rupak Vignesh Swaminathan, Grant Strimel, Jing Liu, and Shinji Watanabe
    In ACLFindings 2026
  37. ASR ACL
    Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
    Zhen Wan, Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye, Ankita Pasad, Szu-Wei Fu, Arushi Goel, Ryo Hachiuma, Shizhe Diao, Kunal Dhawan, Sreyan Ghosh, Yusuke Hirota, Zhehuai Chen, Rafael Valle, Chenhui Chu, Shinji Watanabe, Boris Ginsburg, and Yu-Chiang Frank Wang
    In ACL 2026
  38. Dialogue ACLFindings
    Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
    Siddhant Arora, Jinchuan Tian, Jiatong Shi, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, and Shinji Watanabe
    In ACLFindings 2026
  39. Speech-LLM ICLR
    UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
    Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, and Wei Ping
    In ICLR 2026
  40. SE ICLR
    MAPSS: Manifold-based Assessment of Perceptual Source Separation
    Amir Ivry, Samuele Cornell, and Shinji Watanabe
    In ICLR 2026
  41. SE ICASSP
    ICASSP 2026 URGENT Speech Enhancement Challenge
    Chenda Li, Wei Wang, Marvin Sach, Wangyou Zhang, Kohei Saijo, Samuele Cornell, Yihui Fu, Zhaoheng Ni, Tim Fingscheidt, Shinji Watanabe, and Yanmin Qian
    In ICASSP 2026
  42. ASR ICASSP
    SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
    Pu Wang, Shinji Watanabe, and Hugo Van hamme
    In ICASSP 2026
  43. Speech-LLM ICASSP
    Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
    Bo-Hao Su, Hui-Ying Shih, Jinchuan Tian, Jiatong Shi, Chi-Chun Lee, Carlos Busso, and Shinji Watanabe
    In ICASSP 2026
  44. SE ICASSP
    2025 URGENT Speech Enhancement Challenge Multilingual P.808 Listening Tests: Approach and Results
    Marvin Sach, Yihui Fu, Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar, Wei Wang, Yanmin Qian, Shinji Watanabe, and Tim Fingscheidt
    In ICASSP 2026
  45. Evaluation ICASSP
    Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
    Guan-Ting Lin, Shih-Yun Shan Kuan, Qirui Wang, Jiachen Lian, Tingle Li, Shinji Watanabe, and Hung-yi Lee
    In ICASSP 2026
  46. ASR ICASSP
    CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
    Muhammad Shakeel, Yosuke Fukumoto, Chikara Maeda, Chyi-Jiunn Lin, and Shinji Watanabe
    In ICASSP 2026
  47. Tokenizer ICASSP
    Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means
    Kentaro Onda, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, and Shinji Watanabe
    In ICASSP 2026
  48. SSL ICASSP
    Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context
    Keita Goto, Takashi Maekaku, Jin Sakuma, Jinchuan Tian, Yusuke Shinohara, and Shinji Watanabe
    In ICASSP 2026
  49. Evaluation EACL
    CSPB: Conversational Speech Processing Benchmark for Self-supervised Speech Models
    Zili Huang, Matthew Maciejewski, Leibny Paola Garcia Perera, Shinji Watanabe, and Sanjeev Khudanpur
    In EACL 2026
  50. Tokenizer EACLFindings
    BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
    Haoran Wang, Jiatong Shi, Jinchuan Tian, Bohan Li, Kai Yu, and Shinji Watanabe
    In EACLFindings 2026