Result Details

State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data

BARAHONA, S.; MOSNER, L.; STAFYLAKIS, T.; PLCHOT, O.; PENG, J.; BURGET, L.; CERNOCKY, J. State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data. In Asru 2025 2025 IEEE Automatic Speech Recognition and Understanding Workshop. Honolulu, Hawaii, USA: Institute of Electrical and Electronics Engineers Inc., 2025. p. 1-7. ISBN: 979-8-3315-4426-3.
Type
conference paper
Language
English
Authors
Barahona Sara
Mošner Ladislav, Ing., Ph.D., DCGM (FIT)
Stafylakis Themos
Plchot Oldřich, Ing., Ph.D., DCGM (FIT)
Peng Junyi, DCGM (FIT)
Burget Lukáš, doc. Ing., Ph.D., DCGM (FIT)
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Abstract

In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments be-longing to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation.

Keywords

pretrained models | speaker verification | weak-supervision

URL
Published
2025
Pages
7
Proceedings
Asru 2025 2025 IEEE Automatic Speech Recognition and Understanding Workshop
Conference
2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
ISBN
979-8-3315-4426-3
Publisher
Institute of Electrical and Electronics Engineers Inc.
Place
Honolulu, Hawaii, USA
DOI
EID Scopus
BibTeX
@inproceedings{BUT212163,
  author="{} and Ladislav {Mošner} and  {} and Oldřich {Plchot} and Junyi {Peng} and Lukáš {Burget} and Jan {Černocký}",
  title="State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data",
  booktitle="Asru 2025 2025 IEEE Automatic Speech Recognition and Understanding Workshop",
  year="2025",
  pages="7",
  publisher="Institute of Electrical and Electronics Engineers Inc.",
  address="Honolulu, Hawaii, USA",
  doi="10.1109/ASRU65441.2025.11434721",
  isbn="979-8-3315-4426-3",
  url="https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11434721"
}
Files
Projects
Exchanges for SPEech ReseArch aNd TechnOlogies, EU, Horizon 2020, start: 2021-01-01, end: 2025-12-31, completed
Research groups
Departments
Back to top