Result Details

Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing

PENG, J.; ZHANG, L.; HAN, J.; PLCHOT, O.; ROHDIN, J.; STAFYLAKIS, T.; WANG, S.; ČERNOCKÝ, J. Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing. ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Barcelona, Španělské království: IEEE, 2026. p. 16442.ISBN: 979-8-3315-6701-9.
Type
conference paper
Language
English
Authors
Peng Junyi, DCGM (FIT)
Zhang Lin, Ph.D.
Han Jiangyu, DCGM (FIT)
Plchot Oldřich, Ing., Ph.D., DCGM (FIT)
Rohdin Johan Andréas, M.Sc., Ph.D., FIT (FIT), DCGM (FIT)
Stafylakis Themos, Prof.
Wang Shuai
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Abstract

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typically separate it from taskspecific fine-tuning. This multi-stage approach struggles to create optimal architectures tailored for diverse downstream tasks. In this work, we introduce a unified framework that integrates structured pruning into the downstream fine-tuning process. Our framework unifies these steps, jointly optimizing for task performance and model sparsity in a single stage. This allows the model to learn a compressed architecture specifically for the end task, eliminating the need for complex multi-stage pipelines and knowledge distillation. Our pruned models achieve up to a 70% parameter reduction with negligible performance degradation on large-scale datasets, attaining equal error rates of 0.7%, 0.8%, and 1.6% on Vox1-O, -E, and -H, respectively. Furthermore, our approach demonstrates improved generalization in low-resource scenarios, reducing overfitting and achieving a state-of-the-art 3.7% EER on ASVspoof5.

Keywords

Self-supervised learning, speaker verification, anti-spoofing, structured pruning, fine-tuning

URL
Published
2026
Pages
16442–16446
Proceedings
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Conference
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
ISBN
979-8-3315-6701-9
Publisher
IEEE
Place
Barcelona, Španělské království
DOI
BibTeX
@inproceedings{BUT212034,
  author="Junyi {Peng} and Lin {Zhang} and Jiangyu {Han} and Oldřich {Plchot} and Johan Andréas {Rohdin} and Themos {Stafylakis} and Shuai {Wang} and Jan {Černocký}",
  title="Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing",
  booktitle="ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)",
  year="2026",
  pages="16442--16446",
  publisher="IEEE",
  address="Barcelona, Španělské království",
  doi="10.1109/icassp55912.2026.11464156",
  isbn="979-8-3315-6701-9",
  url="https://ieeexplore.ieee.org/document/11464156"
}
Files
Projects
Linguistics, Artificial Intelligence and Language and Speech Technologies: from Research to Applications, EU, MEZISEKTOROVÁ SPOLUPRÁCE, EH23_020/0008518, start: 2025-01-01, end: 2028-12-31, running
Multilingual and Cross-cultural interactions for context-aware, and bias-controlled dialogue systems for safety-critical applications, EU, HORIZON EUROPE, start: 2024-01-01, end: 2026-12-31, running
Research groups
Departments
Back to top