Result Details
VBx for End-to-End Neural and Clustering-Based Diarization
Han Jiangyu, DCGM (FIT)
Delcroix Marc, FIT (FIT)
Tawara Naohiro
Burget Lukáš, doc. Ing., Ph.D., DCGM (FIT)
We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based EEND model with WavLM features to infer frame-level speaker activity within short windows. The identities and counts of global speakers are then derived in the second stage by clustering speaker embeddings across windows. The focus of this work is to improve the second stage; we filter unreliable embeddings from short segments and reassign them after clustering. We also integrate the VBx clustering to improve robustness when the number of speakers is large and individual speaking durations are limited. Evaluation on a compound benchmark spanning multiple domains is conducted without fine-tuning the EEND model or tuning clustering parameters per dataset. Despite this, the system generalizes well and matches or exceeds recent state-of-the-art performance.
speaker diarization, EEND-VC, VBx, pyannote
@inproceedings{BUT212035,
author="Petr {Pálka} and Jiangyu {Han} and Marc {Delcroix} and {} and Lukáš {Burget}",
title="VBx for End-to-End Neural and Clustering-Based Diarization",
booktitle="ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)",
year="2026",
pages="17472--17476",
publisher="IEEE",
address="Barcelona, Španělské království",
doi="10.1109/icassp55912.2026.11462054",
isbn="979-8-3315-6701-9",
url="https://ieeexplore.ieee.org/document/11462054"
}