2026 Papers
- SS&SD&ASR TASLPOnline Frontend System for Multi-Talker DSR Using Neural Blind Source Separation and DiarizationIEEE/ACM Transactions on Audio, Speech, and Language Processing 2026
- ASR OJSPUncertainty-Based Streaming ASR with Evidential Deep LearningIEEE Open Journal of Signal Processing 2026
- Multimodal TMMTMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different LanguagesIEEE Transactions on Multimedia 2026
- SS&ASR CSLAn End-to-End Integration of Speech Separation and Recognition with Self-Supervised Learning RepresentationComputer Speech & Language 2026
- Evaluation EMNLPBenchmarking Speech-to-Speech Translation ModelsIn Proceedings of Findings of EMNLP 2026
- Speech-LLM EMNLPLong Listening Thoughts: Eliciting Open Auditory Reasoning with Deliberative Perception and Cognitive RefinementIn Proceedings of Findings of EMNLP 2026
- Evaluation EMNLPMP-Bench: Evaluating Voice Agents as a Multiparty Conversation ParticipantIn Proceedings of Findings of EMNLP 2026
- ASR&Speech-LLM EMNLPListen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State InteractionsIn Proceedings of Findings of EMNLP 2026
- Speech-LLM COLMBagpiper: Solving Open-Ended Audio Tasks via Rich CaptionsIn Proceedings of COLM 2026
- ASR InterspeechYODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual SpeechIn Proceedings of Interspeech 2026
- ASR&SD InterspeechMind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker DiarizationIn Proceedings of Interspeech 2026
- Speech-LLM&SD InterspeechGrounding Spoken LLMs in Multi-Speaker Audio via Diarization ConditioningIn Proceedings of Interspeech 2026
- Speech-LLM InterspeechBagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-CaptionIn Proceedings of Interspeech 2026
- TTS&Speech-LLM InterspeechBagpiper-TTS: Natural Language Guided Universal Speech SynthesisIn Proceedings of Interspeech 2026
- Evaluation InterspeechANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality ModelingIn Proceedings of Interspeech 2026
- Speech-LLM InterspeechAn Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and GenerationIn Proceedings of Interspeech 2026
- ASR InterspeechAn Empirical Recipe for Universal Phone RecognitionIn Proceedings of Interspeech 2026
- ASR&Speaker InterspeechSpeaker-Aware Hypothesis Clustering and Merging for Target-Speaker-free and Target-Speaker Multi-Talker ASRIn Proceedings of Interspeech 2026
- SE&Evaluation InterspeechURGENT-MOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality AssessmentIn Proceedings of Interspeech 2026
- SSL InterspeechOnline Predictive Coding for Dual-Mode Self-Supervised Speech ModelsIn Proceedings of Interspeech 2026
- Speech-LLM InterspeechAdapting Text LLMs to Speech via Multimodal Depth Up-ScalingIn Proceedings of Interspeech 2026
- Dialogue InterspeechEndpoint Anticipation for Low-Latency Spoken DialogueIn Proceedings of Interspeech 2026
- ASR InterspeechESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model EraIn Proceedings of Interspeech 2026
- Evaluation InterspeechWho Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware MetricsIn Proceedings of Interspeech 2026
- Evaluation InterspeechEvaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in MeetingsIn Proceedings of Interspeech 2026
- ER&Speech-LLM InterspeechComparative Reasoning: Making an Audio Language Model Better at Comparing EmotionsIn Proceedings of Interspeech 2026
- SSL&Tokenizer InterspeechDissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec TokensIn Proceedings of Interspeech 2026
- ASR InterspeechWhich Data Matter? Embedding-Based Data Selection for Speech RecognitionIn Proceedings of Interspeech 2026
- Speech-LLM ICMLADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools – From Consensus Learning to Ambiguity-Driven Emotion ReasoningIn ICML 2026
- Dialogue ICMLStream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool UsageIn ICML 2026
- Evaluation ICMLLALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken DialoguesIn ICML 2026
- Speech-LLM ICMLAudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion ForcingIn ICML 2026
- Evaluation ACLFull-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated ExaminerIn ACL 2026
- ASR ACLPOWSM: A Phonetic Open Whisper-Style Speech Foundation ModelIn ACL 2026
- Evaluation ACLPRiSM: Benchmarking Phone Realization in Speech ModelsIn ACL 2026
- SLU ACLFindingsPlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio UnderstandingIn ACLFindings 2026
- ASR ACLSpeech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni PerceptionIn ACL 2026
- Dialogue ACLFindingsOptimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI FeedbackIn ACLFindings 2026
- Speech-LLM ICLRUALM: Unified Audio Language Model for Understanding, Generation and ReasoningIn ICLR 2026
- SE ICLRMAPSS: Manifold-based Assessment of Perceptual Source SeparationIn ICLR 2026
- SE ICASSPICASSP 2026 URGENT Speech Enhancement ChallengeIn ICASSP 2026
- ASR ICASSPSSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech RecognitionIn ICASSP 2026
- Speech-LLM ICASSPReasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion RecognitionIn ICASSP 2026
- SE ICASSP2025 URGENT Speech Enhancement Challenge Multilingual P.808 Listening Tests: Approach and ResultsIn ICASSP 2026
- Evaluation ICASSPFull-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech ModelsIn ICASSP 2026
- ASR ICASSPCALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASRIn ICASSP 2026
- Tokenizer ICASSPPhonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-MeansIn ICASSP 2026
- SSL ICASSPOnline Register for Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future ContextIn ICASSP 2026
- Evaluation EACLCSPB: Conversational Speech Processing Benchmark for Self-supervised Speech ModelsIn EACL 2026
- Tokenizer EACLFindingsBSCodec: A Band-Split Neural Codec for High-Quality Universal Audio ReconstructionIn EACLFindings 2026