Result Details
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
Zhang Lin, Ph.D.
Han Jiangyu, DCGM (FIT)
Plchot Oldřich, Ing., Ph.D., DCGM (FIT)
Rohdin Johan Andréas, M.Sc., Ph.D., FIT (FIT), DCGM (FIT)
Stafylakis Themos, Prof.
Wang Shuai
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typically separate it from taskspecific fine-tuning. This multi-stage approach struggles to create optimal architectures tailored for diverse downstream tasks. In this work, we introduce a unified framework that integrates structured pruning into the downstream fine-tuning process. Our framework unifies these steps, jointly optimizing for task performance and model sparsity in a single stage. This allows the model to learn a compressed architecture specifically for the end task, eliminating the need for complex multi-stage pipelines and knowledge distillation. Our pruned models achieve up to a 70% parameter reduction with negligible performance degradation on large-scale datasets, attaining equal error rates of 0.7%, 0.8%, and 1.6% on Vox1-O, -E, and -H, respectively. Furthermore, our approach demonstrates improved generalization in low-resource scenarios, reducing overfitting and achieving a state-of-the-art 3.7% EER on ASVspoof5.
Self-supervised learning, speaker verification, anti-spoofing, structured pruning, fine-tuning
@inproceedings{BUT212034,
author="Junyi {Peng} and Lin {Zhang} and Jiangyu {Han} and Oldřich {Plchot} and Johan Andréas {Rohdin} and Themos {Stafylakis} and Shuai {Wang} and Jan {Černocký}",
title="Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing",
booktitle="ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)",
year="2026",
pages="16442--16446",
publisher="IEEE",
address="Barcelona, Španělské království",
doi="10.1109/icassp55912.2026.11464156",
isbn="979-8-3315-6701-9",
url="https://ieeexplore.ieee.org/document/11464156"
}
Multilingual and Cross-cultural interactions for context-aware, and bias-controlled dialogue systems for safety-critical applications, EU, HORIZON EUROPE, start: 2024-01-01, end: 2026-12-31, running