Result Details

Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams

HE, X.; POLOK, A.; VILLALBA, J.; THEBAUD, T.; MACIEJEWSKI, M. Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams. ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Barcelona, Španělské království: IEEE, 2026. p. 16727.ISBN: 979-8-3315-6701-9.
Type
conference paper
Language
English
Authors
He Xiluo
Polok Alexander, Ing., DCGM (FIT)
Villalba Jesús
Thebaud Thomas
Maciejewski Matthew
Abstract

An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping speech. Although effective, these systems require running the ASR model once per speaker, resulting in inference costs that scale with the number of speakers and limiting their practicality. In this work, we propose a method that decouples the inference cost of activity-conditioned ASR systems from the number of speakers by converting speaker-specific activity outputs into two speaker-agnostic streams. A central challenge is that naïvely merging speaker activities into streams significantly degrades recognition, since pretrained ASR models assume contiguous, single-speaker inputs. To address this, we design new heuristics aimed at preserving conversational continuity and maintaining compatibility with existing systems. We show that our approach is compatible with Diarization-Conditioned Whisper (DiCoW) to greatly reduce runtimes on the AMI and ICSI meeting datasets while retaining competitive performance.

Keywords

Multi-talker ASR, Target-speaker ASR, Whisper, DiCoW

URL
Published
2026
Pages
16727–16731
Proceedings
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Conference
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
ISBN
979-8-3315-6701-9
Publisher
IEEE
Place
Barcelona, Španělské království
DOI
BibTeX
@inproceedings{BUT212031,
  author="{} and Alexander {Polok} and  {} and  {} and  {}",
  title="Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams",
  booktitle="ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)",
  year="2026",
  pages="16727--16731",
  publisher="IEEE",
  address="Barcelona, Španělské království",
  doi="10.1109/icassp55912.2026.11461880",
  isbn="979-8-3315-6701-9",
  url="https://ieeexplore.ieee.org/document/11461880"
}
Files
Projects
Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications, EU, MEZISEKTOROVÁ SPOLUPRÁCE, EH23_020/0008518, start: 2025-01-01, end: 2028-12-31, running
Research groups
Departments
Back to top