Detail výsledku

Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing

PENG, J.; ZHANG, L.; HAN, J.; PLCHOT, O.; ROHDIN, J.; STAFYLAKIS, T.; WANG, S.; ČERNOCKÝ, J. Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing. ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Barcelona, Španělské království: IEEE, 2026. p. 16442.ISBN: 979-8-3315-6701-9.
Typ
článek ve sborníku konference
Jazyk
angličtina
Autoři
Peng Junyi, UPGM (FIT)
Zhang Lin, Ph.D.
Han Jiangyu, UPGM (FIT)
Plchot Oldřich, Ing., Ph.D., UPGM (FIT)
Rohdin Johan Andréas, M.Sc., Ph.D., FIT (FIT), UPGM (FIT)
Stafylakis Themos, Prof.
Wang Shuai
Černocký Jan, prof. Dr. Ing., UPGM (FIT)
Abstrakt

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typically separate it from taskspecific fine-tuning. This multi-stage approach struggles to create optimal architectures tailored for diverse downstream tasks. In this work, we introduce a unified framework that integrates structured pruning into the downstream fine-tuning process. Our framework unifies these steps, jointly optimizing for task performance and model sparsity in a single stage. This allows the model to learn a compressed architecture specifically for the end task, eliminating the need for complex multi-stage pipelines and knowledge distillation. Our pruned models achieve up to a 70% parameter reduction with negligible performance degradation on large-scale datasets, attaining equal error rates of 0.7%, 0.8%, and 1.6% on Vox1-O, -E, and -H, respectively. Furthermore, our approach demonstrates improved generalization in low-resource scenarios, reducing overfitting and achieving a state-of-the-art 3.7% EER on ASVspoof5.

Klíčová slova

Self-supervised learning, speaker verification, anti-spoofing, structured pruning, fine-tuning

URL
Rok
2026
Strany
16442–16446
Sborník
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Konference
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
ISBN
979-8-3315-6701-9
Vydavatel
IEEE
Místo
Barcelona, Španělské království
DOI
BibTeX
@inproceedings{BUT212034,
  author="Junyi {Peng} and Lin {Zhang} and Jiangyu {Han} and Oldřich {Plchot} and Johan Andréas {Rohdin} and Themos {Stafylakis} and Shuai {Wang} and Jan {Černocký}",
  title="Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing",
  booktitle="ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)",
  year="2026",
  pages="16442--16446",
  publisher="IEEE",
  address="Barcelona, Španělské království",
  doi="10.1109/icassp55912.2026.11464156",
  isbn="979-8-3315-6701-9",
  url="https://ieeexplore.ieee.org/document/11464156"
}
Soubory
Projekty
Jazykověda, umělá inteligence a jazykové a řečové technologie: od výzkumu k aplikacím, EU, MEZISEKTOROVÁ SPOLUPRÁCE, EH23_020/0008518, zahájení: 2025-01-01, ukončení: 2028-12-31, řešení
Multilingvální a mezikulturní interakce v dialogových systémech pro bezpečnostně kritické aplikace závislé na kontextu a kontrolou zaujatosti, EU, HORIZON EUROPE, zahájení: 2024-01-01, ukončení: 2026-12-31, řešení
Výzkumné skupiny
Pracoviště
Nahoru