Result Details

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

POLOK, A.; KLEMENT, D.; CORNELL, S.; WIESNER, M.; ČERNOCKÝ, J.; KHUDANPUR, S.; BURGET, L. SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper. ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Barcelona, Španělské království: IEEE, 2026. p. 16712.ISBN: 979-8-3315-6701-9.
Type
conference paper
Language
English
Authors
Polok Alexander, Ing., DCGM (FIT)
Klement Dominik, Ing., DCGM (FIT)
Cornell Samuele
Wiesner Matthew
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Khudanpur Sanjeev
Burget Lukáš, doc. Ing., Ph.D., DCGM (FIT)
Abstract

Speaker-attributed automatic speech recognition (ASR) in multispeaker
environments remains a major challenge. While some
approaches achieve strong performance when fine-tuned on specific
domains, few systems generalize well across out-of-domain datasets.
Our prior work, Diarization-Conditioned Whisper (DiCoW), leverages
speaker diarization outputs as conditioning information and,
with minimal fine-tuning, demonstrated strong multilingual and
multi-domain performance. In this paper, we address a key limitation
of DiCoW: ambiguity in Silence–Target–Non-target–Overlap
(STNO) masks, where two or more fully overlapping speakers may
have nearly identical conditioning despite differing transcriptions.
We introduce SE-DiCoW (Self-Enrolled Diarization-Conditioned
Whisper), which uses diarization output to locate an enrollment
segment anywhere in the conversation where the target speaker is
most active. This enrollment segment is used as fixed conditioning
via cross-attention at each encoder layer. We further refine DiCoW
with improved data segmentation, model initialization, and augmentation.
Together, these advances yield substantial gains: SE-DiCoW
reduces macro-averaged tcpWER by 52.4% relative to the original
DiCoW on the EMMA MT-ASR benchmark.

Keywords

target-speaker ASR, DiCoW, diarization conditioning, multi-speaker ASR, Whisper

URL
Published
2026
Pages
16712–16716
Proceedings
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Conference
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
ISBN
979-8-3315-6701-9
Publisher
IEEE
Place
Barcelona, Španělské království
DOI
BibTeX
@inproceedings{BUT212029,
  author="Alexander {Polok} and Dominik {Klement} and  {} and  {} and Jan {Černocký} and  {} and Lukáš {Burget}",
  title="SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper",
  booktitle="ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)",
  year="2026",
  pages="16712--16716",
  publisher="IEEE",
  address="Barcelona, Španělské království",
  doi="10.1109/icassp55912.2026.11461785",
  isbn="979-8-3315-6701-9",
  url="https://ieeexplore.ieee.org/document/11461785"
}
Files
Projects
Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications, EU, MEZISEKTOROVÁ SPOLUPRÁCE, EH23_020/0008518, start: 2025-01-01, end: 2028-12-31, running
Research groups
Departments
Back to top