Result Details

Efficient and Robust Speaker Diarization via Structured Pruning of Self-Supervised Models

HAN, J.; PALKA, P.; DELCROIX, M.; LANDINI, F.; ROHDIN, J.; ČERNOCKÝ, J.; BURGET, L. Efficient and Robust Speaker Diarization via Structured Pruning of Self-Supervised Models. IEEE transactions on audio, speech, and language processing, 2026, vol. 34, iss. 2026, p. 1903-1914.
Type
journal article
Language
English
Authors
Abstract

This work presents a framework for compressing self-supervised models for speaker diarization through structured pruning guided by knowledge distillation. We investigate pruning objectives that target reducing both model parameters and computational complexity, where knowledge distillation enables compact models and pruning removes unnecessary parameters to improve hardware efficiency. We further analyze alternative pruning strategies, showing that a simple overall pruning approach provides the best balance between efficiency and accuracy. Compared to the original unpruned model, our method achieves up to 80% model size reduction and 4x faster inference without performance degradation. Comprehensive experiments across eight public diarization datasets demonstrate that the pruned models consistently match or surpass the performance of their uncompressed counterparts. Furthermore, we show strong out-of-domain generalization on the CHiME-6 dataset, achieving accuracy comparable to the top systems in the CHiME-7 challenge without any domain adaptation. These results highlight that structured pruning, when guided by distillation, can yield efficient and generalizable diarization systems suitable for real-world applications.

Keywords

Computational modeling, Adaptation models, Logic gates, Transformers, Training, Stochastic processes, Speech processing, Pipelines, Optimization, Kernel, Speaker diarization, WavLM, model compression, knowledge distillation, structured pruning

URL
Published
2026
Pages
1903–1914
Journal
IEEE transactions on audio, speech, and language processing, vol. 34, no. 2026, ISSN
Publisher
IEEE
DOI
UT WoS
001735991600001
EID Scopus
BibTeX
@article{BUT212048,
  author="Jiangyu {Han} and Petr {Pálka} and Marc {Delcroix} and Federico Nicolás {Landini} and Johan Andréas {Rohdin} and Jan {Černocký} and Lukáš {Burget}",
  title="Efficient and Robust Speaker Diarization via Structured Pruning of Self-Supervised Models",
  journal="IEEE transactions on audio, speech, and language processing",
  year="2026",
  volume="34",
  number="2026",
  pages="1903--1914",
  doi="10.1109/TASLPRO.2026.3675801",
  url="https://ieeexplore.ieee.org/document/11447368"
}
Files
Projects
Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications, EU, MEZISEKTOROVÁ SPOLUPRÁCE, EH23_020/0008518, start: 2025-01-01, end: 2028-12-31, running
Research groups
Departments
Back to top