Members
The publication list under each member is fetched automatically from Google Scholar and shows only that member's most recent papers, so it may be incomplete or out of date, and titles and venues are not always correct. For an accurate and complete list, please see each member's own website or Google Scholar profile, linked on their card. The lab's curated list of papers is on the Publications page.
Faculty

Shinji Watanabe
Publications (842)
- Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models ICASSP, 2026
- AHa-bench: Benchmarking audio hallucinations in large audio-language models Advances in Neural Information Processing Systems 38, 2026
- Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner ACL, 2026
- Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage ICML, 2026
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning ICLR, 2026
- POWSM: A Phonetic Open Whisper-Style Speech Foundation Model ACL, 2026
- Low-resource audio codec (lrac): 2025 challenge description ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and …, 2026
- Tmt: Tri-modal translation between speech, image, and text by processing different modalities as different languages IEEE Transactions on Multimedia, 2026
- ICASSP 2026 URGENT Speech Enhancement Challenge ICASSP, 2026
- Bagpiper: Solving open-ended audio tasks via rich captions arXiv preprint arXiv:2602.05220, 2026
- An end-to-end integration of speech separation and recognition with self-supervised learning representation Computer Speech & Language 95, 101813, 2026
- Music Arena: Live evaluation for text-to-music Advances in Neural Information Processing Systems 38, 2026
- Urgentmos: Unified multi-metric and preference learning for robust speech quality assessment arXiv preprint arXiv:2601.18438, 2026
- Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback ACLFindings, 2026
- PRiSM: Benchmarking Phone Realization in Speech Models ACL, 2026
- Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception ACL, 2026
- Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition ICASSP, 2026
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing ICML, 2026
- Mind the gap: Impact of synthetic conversational data on multi-talker ASR and speaker diarization arXiv preprint arXiv:2605.15442, 2026
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction EACLFindings, 2026
Showing 20 of 842. View all on Google Scholar ↗
Post-Doc

Samuele Cornell
Publications (102)
- ICASSP 2026 URGENT Speech Enhancement Challenge ICASSP, 2026
- Se-dicow: Self-enrolled diarization-conditioned whisper ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and …, 2026
- An end-to-end integration of speech separation and recognition with self-supervised learning representation Computer Speech & Language 95, 101813, 2026
- Urgentmos: Unified multi-metric and preference learning for robust speech quality assessment arXiv preprint arXiv:2601.18438, 2026
- Mind the gap: Impact of synthetic conversational data on multi-talker ASR and speaker diarization arXiv preprint arXiv:2605.15442, 2026
- Dissecting sensitivity to training language in self-supervised speech learning using neural audio codec tokens arXiv preprint arXiv:2607.26350, 2026
- Cross-Talk Speech Reduction, by Separation, for Separation arXiv preprint arXiv:2605.19695, 2026
- 2025 URGENT Speech Enhancement Challenge Multilingual Ṗ808 Listening Tests: Approach and Results ICASSP, 2026
- Ring Mixing with Auxiliary Signal-to-Consistency-Error Ratio Loss for Unsupervised Denoising in Speech Separation arXiv preprint arXiv:2604.08415, 2026
- The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge arXiv preprint arXiv:2601.16273, 2026
- PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding ACLFindings, 2026
- ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era arXiv preprint arXiv:2606.21854, 2026
- Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning arXiv preprint arXiv:2606.18134, 2026
- Exploiting Noise Inseparability for Weakly-Supervised Discriminative Speech Denoising Using Noisy Targets arXiv preprint arXiv:2606.02327, 2026
- MAPSS: Manifold-based Assessment of Perceptual Source Separation ICLR, 2026
- Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics arXiv preprint arXiv:2603.22709, 2026
- Modeling Overlapped Speech with Shuffles arXiv preprint arXiv:2603.17769, 2026
- Interspeech 2025 URGENT Speech Enhancement Challenge Interspeech, 2025
- Lessons Learned from the URGENT 2024 Speech Enhancement Challenge Interspeech, 2025
- ESPnet-SpeechLM: An Open Speech Language Model Toolkit NAACL, 2025
Showing 20 of 102. View all on Google Scholar ↗
PhD Students

Jinchuan Tian
Publications (63)
- Omnivinci: Enhancing architecture and data for omni-modal understanding llm International Conference on Learning Representations 2026, 56101-56138, 2026
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning ICLR, 2026
- Bagpiper: Solving open-ended audio tasks via rich captions arXiv preprint arXiv:2602.05220, 2026
- Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback ACLFindings, 2026
- Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception ACL, 2026
- Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition ICASSP, 2026
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction EACLFindings, 2026
- Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context ICASSP, 2026
- Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks arXiv preprint arXiv:2601.12205, 2026
- An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation arXiv preprint arXiv:2607.02119, 2026
- Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis arXiv preprint arXiv:2606.22811, 2026
- ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era arXiv preprint arXiv:2606.21854, 2026
- Online Predictive Coding for Dual-Mode Self-Supervised Speech Model arXiv preprint arXiv:2606.21268, 2026
- Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption arXiv preprint arXiv:2606.21227, 2026
- VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music NAACL, 2025
- Spoofceleb: Speech deepfake detection and sasv in the wild IEEE Open Journal of Signal Processing 6, 68-77, 2025
- OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Interspeech, 2025
- Preference Alignment Improves Language Model-Based TTS ICASSP, 2025
- OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models ICML, 2025
- OpusLM: A Family of Open Unified Speech Language Models Interspeech, 2025
Showing 20 of 63. View all on Google Scholar ↗

William Chen
Publications (78)
- Steerable vision-language-action policies for embodied reasoning and hierarchical control arXiv preprint arXiv:2602.13193, 2026
- Bagpiper: Solving open-ended audio tasks via rich captions arXiv preprint arXiv:2602.05220, 2026
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing ICML, 2026
- Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control? arXiv preprint arXiv:2608.02547, 2026
- Dissecting sensitivity to training language in self-supervised speech learning using neural audio codec tokens arXiv preprint arXiv:2607.26350, 2026
- Improving Robotic Generalist Policies via Flow Reversal Steering arXiv preprint arXiv:2606.13675, 2026
- An Empirical Recipe for Universal Phone Recognition arXiv preprint arXiv:2603.29042, 2026
- Adapting Generalist Robot Policies with Semantic Reinforcement Learning arXiv preprint arXiv:2606.31958, 2026
- ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era arXiv preprint arXiv:2606.21854, 2026
- Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption arXiv preprint arXiv:2606.21227, 2026
- Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings arXiv preprint arXiv:2606.17542, 2026
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks ICLR, 2025
- Quantum-enabled microwave-to-optical transduction via silicon nanomechanics Nature nanotechnology 20 (5), 602-608, 2025
- OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Interspeech, 2025
- OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models ICML, 2025
- Polaris: Scalable real-to-sim evaluations for generalist robot policies arXiv preprint arXiv:2512.16881, 2025
- Vulbinllm: Llm-powered vulnerability detection for stripped binaries arXiv preprint arXiv:2505.22010, 2025
- OpusLM: A Family of Open Unified Speech Language Models Interspeech, 2025
- Measuring general intelligence with generated games arXiv preprint arXiv:2505.07215, 2025
- Findings of the IWSLT 2025 evaluation campaign PROCEEDINGS OF THE 22ND INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE …, 2025
Showing 20 of 78. View all on Google Scholar ↗

Shikhar Bharadwaj (co-supervising)
Publications (26)
- POWSM: A Phonetic Open Whisper-Style Speech Foundation Model ACL, 2026
- PRiSM: Benchmarking Phone Realization in Speech Models ACL, 2026
- An Empirical Recipe for Universal Phone Recognition arXiv preprint arXiv:2603.29042, 2026
- The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge arXiv preprint arXiv:2601.16273, 2026
- Phone Segmentation and Recognition through Phonological Activation Mapping arXiv preprint arXiv:2607.09020, 2026
- Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities arXiv preprint arXiv:2507.06261, 2025
- OpusLM: A Family of Open Unified Speech Language Models Interspeech, 2025
- ESPnet-SpeechLM: An Open Speech Language Model Toolkit NAACL, 2025
- ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems NAACL, 2025
- OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder WASPAA, 2025
- Context-driven dynamic pruning for large speech foundation models arXiv preprint arXiv:2505.18860, 2025
- Identifying and Mitigating Mismatched Language Code in Multilingual ASR ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and …, 2025
- Identifying and Mitigating Mismatched Language Signal in Multilingual Automated Speech Recognition US Patent App. 19/173,179, 2025
- EmoNews: A Spoken Dialogue System for Expressive News Conversations Proceedings of the 26th Annual Meeting of the Special Interest Group on …, 2025
- VERSA-v2: A Modular and Scalable Toolkit for Speech and Audio Evaluation with Expanded Metrics, Visualization, and LLM Integration ASRU, 2025
- Indicgenbench: A multilingual benchmark to evaluate generation capabilities of llms on indic languages Proceedings of the 62nd Annual Meeting of the Association for Computational …, 2024
- Codequeries: A dataset of semantic queries over code Proceedings of the 17th Innovations in Software Engineering Conference, 1-11, 2024
- STAB: speech tokenizer assessment benchmark arXiv preprint arXiv:2409.02384, 2024
- Multimodal modeling for spoken language identification ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and …, 2024
- Label aware speech representation learning for language identification arXiv preprint arXiv:2306.04374, 2023
Showing 20 of 26. View all on Google Scholar ↗

Chien-yu Huang
Publications (16)
- Desta2. 5-audio: Toward general-purpose large audio language model with self-generated cross-modal alignment IEEE Transactions on Audio, Speech and Language Processing, 2026
- Bagpiper: Solving open-ended audio tasks via rich captions arXiv preprint arXiv:2602.05220, 2026
- Causal tracing of audio-text fusion in large audio language models arXiv preprint arXiv:2603.13768, 2026
- PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding ACLFindings, 2026
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks ICLR, 2025
- A preliminary exploration with gpt-4o voice mode arXiv preprint arXiv:2502.09940, 2025
- SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and …, 2025
- Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech ICASSP, 2024
- Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition SLT, 2024
- Prompting and adapter tuning for self-supervised encoder-decoder speech model 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 1-8, 2023
- Toward degradation-robust voice conversion ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and …, 2022
- Investigating on incorporating pretrained and learnable speaker representations for multi-speaker multi-style text-to-speech ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and …, 2021
- Defending your voice: Adversarial attack on voice conversion 2021 IEEE Spoken Language Technology Workshop (SLT), 552-559, 2021
- Utilizing self-supervised representations for MOS prediction Interspeech 2021, 2021
- How far are we from robust voice conversion: A survey 2021 IEEE Spoken Language Technology Workshop (SLT), 514-521, 2021
- Improving cross-lingual reading comprehension with self-training arXiv preprint arXiv:2105.03627, 2021

Jaeyeon Kim (co-supervising)
Publications (13)
- The interspeech 2026 audio reasoning challenge: Evaluating reasoning process quality for audio reasoning models and agents arXiv preprint arXiv:2602.14224, 2026
- Wow-bench: Evaluating fine-grained acoustic perception in audio-language models via marine mammal vocalizations Findings of the Association for Computational Linguistics: ACL 2026, 31208-31229, 2026
- K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts arXiv preprint arXiv:2606.02404, 2026
- Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions arXiv preprint arXiv:2606.24082, 2026
- Towards Scene-Aware Video-to-Spatial Audio Generation International Journal of Computer Vision 134 (4), 184, 2026
- Visage: Video-to-spatial audio generation ICLR 2025, 2025
- Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning ICASSP 2026, 2025
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span NeurIPS 2025, 2025
- Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning ICASSP 2024, 2024
- Enclap++: Analyzing the enclap framework for optimizing automated audio captioning performance DCASE2024 Workshop, 2024
- Expanding on EnCLAP with auxiliary retrieval model for automated audio captioning DCASE2024 Challenge Technical Report, 2024
- Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations ICASSP 2024, 2024
- PITS: Variational pitch inference without fundamental frequency for end-to-end pitch-controllable TTS ICML 2023 Workshop on SPIGM, 2023

Sichen Jin
Publications (13)
- Electronic device and method for voice recognition US Patent App. 18/828,519, 2025
- Electronic apparatus for speech recognition, and controlling method thereof US Patent 12,205,576, 2025
- CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting Interspeech, 2958-1796, 2024
- Electronic apparatus and control method thereof US Patent 11,893,980, 2024
- Electronic device and method for controlling the same US Patent App. 18/057,491, 2023
- Conformer-based on-device streaming speech recognition with KD compression and two-pass architecture 2022 IEEE Spoken Language Technology Workshop (SLT), 92-99, 2023
- A More Accurate Internal Language Model Score Estimation for the Hybrid Autoregressive Transducer Proc. Interspeech 2023, 869-873, 2023
- Server that supports speech recognition of device, and operation method of the server US Patent 11,514,916, 2022
- Electronic apparatus, control method thereof and electronic system US Patent App. 17/574,214, 2022
- Streaming on-device end-to-end asr system for privacy-sensitive voice-typing Proc. Interspeech 2020, 3371-3375, 2020
- Utterance Invariant Training for Hybrid Two-Pass End-to-End Speech Recognition. Interspeech, 2827-2831, 2020
- Attention based on-device streaming speech recognition with large speech corpus 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU …, 2019
- Less is Enough: A Target Fine-tuning Strategy for Contextual Biasing Speech Recognition

Masao Someki
Publications (12)
- Bagpiper: Solving open-ended audio tasks via rich captions arXiv preprint arXiv:2602.05220, 2026
- PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding ACLFindings, 2026
- ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era arXiv preprint arXiv:2606.21854, 2026
- Towards inclusive ASR: investigating voice conversion for dysarthric speech recognition in low-resource languages arXiv preprint arXiv:2505.14874, 2025
- Context-aware Dynamic Pruning for Speech Foundation Models ICLR, 2025
- On-device Streaming Discrete Speech Units Interspeech, 2025
- Context-driven dynamic pruning for large speech foundation models arXiv preprint arXiv:2505.18860, 2025
- Singingsds: A singing-capable spoken dialogue system for conversational roleplay applications arXiv preprint arXiv:2511.20972, 2025
- ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration SLT, 2024
- Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference ASRU, 2023
- Espnet-onnx: Bridging a gap between research and production 2022 Asia-Pacific Signal and Information Processing Association Annual …, 2022
- A comparative study on transformer vs rnn in speech applications ASRU, 2019

Chin-Jou Li (co-supervising)
Publications (11)
- POWSM: A Phonetic Open Whisper-Style Speech Foundation Model ACL, 2026
- Prompt-MII: Meta-learning instruction induction for LLMs International Conference on Learning Representations 2026, 52735-52756, 2026
- Bagpiper: Solving open-ended audio tasks via rich captions arXiv preprint arXiv:2602.05220, 2026
- PRiSM: Benchmarking Phone Realization in Speech Models ACL, 2026
- Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease arXiv preprint arXiv:2603.22225, 2026
- Efficient many-shot in-context learning with dynamic block-sparse attention Proceedings of the 63rd Annual Meeting of the Association for Computational …, 2025
- Towards inclusive ASR: investigating voice conversion for dysarthric speech recognition in low-resource languages arXiv preprint arXiv:2505.14874, 2025
- Face swapping in seizure videos for patient deidentification Epilepsy Research 207, 107453, 2024
- Epileptic Seizure Classification with Patient-level and Video-level Contrastive Pretraining 2024 46th Annual International Conference of the IEEE Engineering in …, 2024
- Deep complex u-net with conformer for audio-visual speech enhancement arXiv preprint arXiv:2309.11059, 2023
- Artificial intelligence-based face transformation in patient seizure videos for privacy protection Mayo Clinic Proceedings: Digital Health 1 (4), 619-628, 2023

Ting (Justin) Jiang
Publications (6)
- ZEUS: Accelerating Diffusion Models with Only Second-Order Predictor ACM MM 2026 (Oral), 2026
- T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning arXiv preprint arXiv:2603.03790, 2026
- DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions arXiv preprint arXiv:2607.20469, 2026
- SADA: Stability-guided Adaptive Diffusion Acceleration ICML 2025, 2025
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models AAAI 2026, 2025
- Enhanced Cyclic Coordinate Descent Methods for Elastic Net Penalized Linear Models NeurIPS 2025, 2025

Haeri Kim (co-supervising)
Publications (3)
- BBPE16: UTF-16-based byte-level byte-pair encoding for improved multilingual speech recognition ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and …, 2026
- A More Accurate Internal Language Model Score Estimation for the Hybrid Autoregressive Transducer Proc. Interspeech 2023, 869-873, 2023
- Self-diagnosing gan: Diagnosing underrepresented samples in generative adversarial networks Advances in Neural Information Processing Systems 34, 1925-1938, 2021
Master Students

Haoran Wang
Publications (11)
- Bagpiper: Solving open-ended audio tasks via rich captions COLM 2026, 2026
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction EACLFindings, 2026
- An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation Interspeech 2026, 2026
- Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis Interspeech 2026, 2026
- Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption Interspeech 2026, 2026
- A Unified and Reproducible Experimentation Framework for Speech Understanding Interspeech 2026, 2026
- Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy ICASSP 2026, 2025
- PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning ASRU, 2025
- Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency ASRU 2025, 2025
- VERSA-v2: A Modular and Scalable Toolkit for Speech and Audio Evaluation with Expanded Metrics, Visualization, and LLM Integration ASRU, 2025
- Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective ICASSP 2026, 2024
Visitors

Thanapat Trachu
Publications (7)
- Quantifying speaker embedding phonological rule interactions in accented speech synthesis ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and …, 2026
- Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data arXiv preprint arXiv:2603.07534, 2026
- Auditory-Inspired Transformer for Binaural Speech Enhancement and Spatial Cue Preservation ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and …, 2026
- Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech arXiv preprint arXiv:2603.07551, 2026
- Decom-renorm-merge: Model merging on the right space improves multitasking arXiv preprint arXiv:2505.23117, 2025
- Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing arXiv preprint arXiv:2506.11542, 2025
- Thunder: Unified regression-diffusion speech enhancement with a single reverse step using brownian bridge arXiv preprint arXiv:2406.06139, 2024

Kuang-Da Wang
Publications (31)
- Test-time alignment for large language models via textual model predictive control International Conference on Learning Representations 2026, 317-346, 2026
- Webgen-v bench: Structured representation for enhancing visual design in llm-based web generation and evaluation Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and …, 2026
- Benchmarking Agentic Newswriting via Journalistic Workflows Findings of the Association for Computational Linguistics: ACL 2026, 36450-36463, 2026
- Visual Prompt Discovery via Semantic Exploration arXiv preprint arXiv:2603.16250, 2026
- Adapting to Evolving Data: Test-Time Expert Aggregation for Imbalanced Tabular Regression Proceedings of the Nineteenth ACM International Conference on Web Search and …, 2026
- Multi-Agent Reinforcement Learning Correctable Strategy: A Framework with Correctable Strategies for Portfolio Management Engineering Proceedings 120 (1), 11, 2026
- Template-based financial report generation in agentic and decomposed information retrieval Proceedings of the 48th International ACM SIGIR Conference on Research and …, 2025
- Extending automatic machine translation evaluation to book-length documents Proceedings of the 2025 Conference on Empirical Methods in Natural Language …, 2025
- Mixture experts with test-time self-supervised aggregation for tabular imbalanced regression arXiv preprint arXiv:2506.07033, 2025
- Tree-of-report: Table-to-text generation for sports game reports with tree-structured prompting ACL 2025 Student Research Workshop, 2025
- NEWSAGENT: Benchmarking Multimodal Agents as Journalists with Real-World Newswriting Tasks arXiv preprint arXiv:2509.00446, 2025
- NVIDIA-NeMo’s WMT 2025 Metrics Shared Task Submission Proceedings of the Tenth Conference on Machine Translation, 920-925, 2025
- Ddot: A derivative-directed dual-decoder ordinary differential equation transformer for dynamic system modeling Pacific-Asia Conference on Knowledge Discovery and Data Mining, 434-445, 2025
- APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning Proceedings of the AAAI Conference on Artificial Intelligence 39 (20), 21536 …, 2025
- RallyDiffuser: A Representation-Guided Diffusion Model Framework for Strategic Planning in Badminton Proceedings of the 24th International Conference on Autonomous Agents and …, 2025
- Imitation Learning of Correlated Policies in Stackelberg Games arXiv preprint arXiv:2503.08883, 2025
- Plan2Align: Predictive Planning Based Test-Time Preference Alignment in Paragraph-Level Machine Translation. CoRR, 2025
- Root Cause Analysis In Microservice Using Neural Granger Causal Discovery Proceedings of the AAAI Conference on Artificial Intelligence 38 (1), 206-213, 2024
- BADGE: BADminton report Generation and Evaluation with LLM arXiv preprint arXiv:2406.18116, 2024
- The coachai badminton environment: Bridging the gap between a reinforcement learning environment and real-world badminton games Proceedings of the AAAI Conference on Artificial Intelligence 38 (21), 23844 …, 2024
Showing 20 of 31. View all on Google Scholar ↗
Industrial Collaborators

Jee-weon Jung
Publications (97)
- ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech Computer Speech & Language 95, 101825, 2026
- WildSpoof: advancing in-the-wild data in Text-to-Speech generation and Spoofing-aware automatic speaker verification ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and …, 2026
- Which Data Matter? Embedding-Based Data Selection for Speech Recognition arXiv preprint arXiv:2603.05819, 2026
- Spoofceleb: Speech deepfake detection and sasv in the wild IEEE Open Journal of Signal Processing 6, 68-77, 2025
- X3A: Efficient Multimodal Deepfake Detection with Score-Level Fusion ACM/SIGAPP Symposium on Applied Computing, 767-774, 2025
- Chain-of-Thought Training for Open E2E Spoken Dialogue Systems ISCA Interspeech 2025, 2025
- WildSpoof Challenge Evaluation Plan arXiv preprint arXiv:2508.16858, 2025
- Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels ICASSP, 2025
- SEED: Speaker Embedding Enhancement Diffusion Model ISCA Interspeech 2025, 2025
- Context-Driven Dynamic Pruning for Large Speech Foundation Models ISCA Interspeech 2025, 2025
- Token-based Attractors and Cross-attention in Spoof Diarization IEEE ASRU, 2025
- ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale ASVspoof 2024 Workshop (ISCA Interspeech 2024 Satellite), 2024
- OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer Interspeech, 2024
- Voxtlm: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks ICASSP, 2024
- Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study ICASSP, 2024
- ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models Interspeech, 2024
- ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech SLT, 2024
- One Model to Rule Them All? Towards End-to-End Joint Speaker Diarization and Speech Recognition ICASSP, 2024
- Improving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up Augmentation ICASSP, 2024
- The VoxCeleb Speaker Recognition Challenge: A Retrospective IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024
Showing 20 of 97. View all on Google Scholar ↗
Alumni
Visting Faculty
- 2023. 09 -- 2024. 06: Karen Livescu (TTIC)
Post-Docs
- 2024. 02 - 2025. 05: Hye-jin Shim (CMU)
- 2023. 03 - 2024. 09: Jeeweon Jung (CMU)
- 2022. 03 - 2024. 08: Soumi Maiti (CMU)
- 2021. 09 - 2024. 07: Zhong-Qiu Wang (CMU)
PhD
- 2021. 09 - 2026. 04: Brian Yan (CMU)
- 2022. 05 - 2026. 03: Li-Wei Chen (CMU, co-supervising)
- 2020. 08 - 2026. 03: Siddhant Arora (CMU)
- 2019. 09 - 2025. 12: Jiatong Shi (CMU)
- 2020. 09 - 2025. 04: Yifan Peng (CMU)
- 2019. 09 - 2024. 05: Xuankai Chang (CMU)
- 2021. 09 - 2024. 05: Muqiao Yang (CMU, co-supervisor)
- 2021. 01 - 2023. 06: Xinjian Li (CMU, co-supervisor)
- 2020. 09 - 2023. 06: Jessica Huynh (CMU, co-supervisor)
- 2020. 09 - 2022. 08: Siddharth Dalmia (CMU, co-supervisor)
- 2017. 09 - 2021. 08: Aswin Shanmugam Subramanian (JHU)
- 2017. 10 - 2021. 08: Matthew Maciejewski (JHU, co-supervisor)
- 2017. 12 - 2020. 12: Matthew Wiesner (JHU, co-supervising)
MS & Undergraduate
- 2024. 08 - 2026. 05: Chyi-Jiunn Lin (CMU)
- 2023. 05 - 2025. 05: Kwanghee Choi (CMU)
- 2022. 09 - 2024. 05: Shih-Lun Wu (CMU)
- 2021. 09 - 2023. 05: Dan Berrebbi (CMU)
- 2021. 09 - 2022. 12: Dorsa Zeinali (CMU)
- 2021. 09 - 2022. 12: Karthik Ganesan (MIIS directed study, CMU)
- 2021. 01 - 2022. 08: Chaitanya Narisetty (CMU)
- 2021. 01 - 2022. 08: Peter Wu (CMU, co-supervisor)
- 2021. 09 - 2022. 08: Sujay Suresh Kumar (MIIS directed study, CMU)
- 2021. 09 - 2022. 08: Debayan Ghosh (CMU)
- 2020. 08 - 2021. 08: Tianzi Wang (JHU)
- 2018. 07 - 2019. 06: Zhiqi Wang (JHU)
- 2017. 09 - 2018. 12: Szu-Jui Chen (JHU)
Visitors & Collaborators
- 2026. 02 - 2026. 07: Dahee Yang (Hanyang University)
- 2026. 01 - 2026. 05: Alexander Polok (Brno University of Technology)
- 2025. 12 - 2026. 03: Xun Gong (Shanghai Jiao Tong University)
- 2025. 08 - 2025. 12: Haoran Wang (Shanghai Jiao Tong University)
- 2025. 04 - 2025. 12: Bo-Hao Su (National Tsing Hua University)
- 2025. 05 - 2025. 11: Ji-Hoon Kim (Korea Advanced Institute of Science and Technology)
- 2025. 02 - 2025. 07: Jialu Li (University of Illinois Urbana-Champaign)
- 2025. 01 - 2025. 04: Pu Wang (KU Leuven)
- 2024. 08 - 2025. 02: Kalvin Chang (UC Berkeley)
- 2024. 09 - 2025. 02: Holger Severin Bovbjerg (Aalborg University)
- 2024. 11 - 2024. 12: Junyi Peng (Brno University of Technology)
- 2024. 08 - 2024. 12: Carlos Carvalho (Instituto Superior Técnico )
- 2024. 04 - 2024. 12: Shuichiro Shimizu (Kyoto University)
- 2024. 08 - 2024. 11: Shih-Heng Wang (National Taiwan University)
- 2024. 08 - 2024. 11: Yoshiaki Bando (National Institute of Advanced Industrial Science and Technology)
- 2023. 09 - 2024. 08: Yihan Wu (Renmin University)
- 2023. 11 - 2024. 04: Chenda Li (Shanghai Jiaotong University)
- 2023. 08 - 2024. 03: Roshan Sharma (Carnegie Mellon University)
- 2023. 03 - 2024. 03: Wangyou Zhang (Shanghai Jiaotong University)
- 2023. 08 - 2023. 10: Minsu Kim (Korea Advanced Institute of Science and Technology)
- 2023. 05 - 2023. 08: Kohei Saijo (Waseda University)
- 2022. 10 - 2023. 01: Takaaki Saeki (University of Tokyo)
- 2021. 12 - 2022. 12: Yosuke Kashiwagi (Sony)
- 2022. 04 - 2022. 09: Samuele Cornell (Universita Politecnica delle Marche)
- 2022. 03 - 2022. 06: Yoshiki Masuyama (Tokyo Metropolitan University)
- 2021. 07 - 2022. 07: Yushi Ueda (Japan Patent Office)
- 2021. 08 - 2021. 12: Yen-Ju Lu (Research Center for Information Technology Innovation, Academia Sinica)
- 2020. 01 - 2021. 01: Pengcheng Guo (Northwestern Polytechnical University)
- 2019. 12 - 2020. 03 & 2022. 03 - 2022. 06: Yosuke Higuchi (Waseda University)
- 2019. 12 - 2020. 12: Jing Shi (Chinese Academy of Science)
- 2019. 07 - 2019. 10: Katsuki Inoue (Okayama University)
- 2018. 11 - 2019. 05: Murali Karthick Baskar (Brno University of Technology)
- 2018. 09 - 2018. 12: Xuankai Chang (Shanghai Jiao Tong University)
- 2018. 08 - 2018. 09: Hirofumi Inaguma (Kyoto University)
- 2018. 07 - 2020. 03: Yusuke Fujita (Hitachi Ltd.)
- 2018. 04 - 2018. 09: Nelson Enrique Yalta Soplin (Waseda University)