Result Details
Efficient and Robust Speaker Diarization via Structured Pruning of Self-Supervised Models
Pálka Petr, Ing., DCGM (FIT)
Delcroix Marc, FIT (FIT)
Landini Federico Nicolás, Ph.D.
Rohdin Johan Andréas, M.Sc., Ph.D., FIT (FIT), DCGM (FIT)
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Burget Lukáš, doc. Ing., Ph.D., DCGM (FIT)
This work presents a framework for compressing self-supervised models for speaker diarization through structured pruning guided by knowledge distillation. We investigate pruning objectives that target reducing both model parameters and computational complexity, where knowledge distillation enables compact models and pruning removes unnecessary parameters to improve hardware efficiency. We further analyze alternative pruning strategies, showing that a simple overall pruning approach provides the best balance between efficiency and accuracy. Compared to the original unpruned model, our method achieves up to 80% model size reduction and 4x faster inference without performance degradation. Comprehensive experiments across eight public diarization datasets demonstrate that the pruned models consistently match or surpass the performance of their uncompressed counterparts. Furthermore, we show strong out-of-domain generalization on the CHiME-6 dataset, achieving accuracy comparable to the top systems in the CHiME-7 challenge without any domain adaptation. These results highlight that structured pruning, when guided by distillation, can yield efficient and generalizable diarization systems suitable for real-world applications.
Computational modeling, Adaptation models, Logic gates, Transformers, Training, Stochastic processes, Speech processing, Pipelines, Optimization, Kernel, Speaker diarization, WavLM, model compression, knowledge distillation, structured pruning
@article{BUT212048,
author="Jiangyu {Han} and Petr {Pálka} and Marc {Delcroix} and Federico Nicolás {Landini} and Johan Andréas {Rohdin} and Jan {Černocký} and Lukáš {Burget}",
title="Efficient and Robust Speaker Diarization via Structured Pruning of Self-Supervised Models",
journal="IEEE transactions on audio, speech, and language processing",
year="2026",
volume="34",
number="2026",
pages="1903--1914",
doi="10.1109/TASLPRO.2026.3675801",
url="https://ieeexplore.ieee.org/document/11447368"
}