Detail výsledku

Self-supervised speaker embeddings

STAFYLAKIS, T.; ROHDIN, J.; PLCHOT, O.; MIZERA, P.; BURGET, L. Self-supervised speaker embeddings. In Proceedings of Interspeech. Proceedings of Interspeech. Graz: International Speech Communication Association, 2019. no. 9, p. 2863-2867. ISSN: 1990-9772.

Typ

článek ve sborníku konference

Jazyk

anglicky

Autoři

Stafylakis Themos
Rohdin Johan Andréas, M.Sc., Ph.D., FIT (FIT), UPGM (FIT)
Plchot Oldřich, Ing., Ph.D., UPGM (FIT)
MIZERA, P.
Burget Lukáš, doc. Ing., Ph.D., UPGM (FIT)

Abstrakt

Contrary to i-vectors, speaker embeddings such as x-vectors areincapable of leveraging unlabelled utterances, due to the classificationloss over training speakers. In this paper, we explorean alternative training strategy to enable the use of unlabelledutterances in training. We propose to train speaker embeddingextractors via reconstructing the frames of a target speech segment,given the inferred embedding of another speech segmentof the same utterance. We do this by attaching to the standardspeaker embedding extractor a decoder network, which we feednot merely with the speaker embedding, but also with the estimatedphone sequence of the target frame sequence.The reconstruction loss can be used either as a single objective,or be combined with the standard speaker classificationloss. In the latter case, it acts as a regularizer, encouraging generalizabilityto speakers unseen during training. In all cases, theproposed architectures are trained from scratch and in an endto-end fashion. We demonstrate the benefits from the proposedapproach on the VoxCeleb and Speakers in the Wild Databases,and we report notable improvements over the baseline.

Klíčová slova

speaker recognition, self-supervised learning,deep learning

URL

Rok

2019

Strany

2863–2867

Časopis

Proceedings of Interspeech, roč. 2019, č. 9, ISSN 1990-9772

Sborník

Proceedings of Interspeech

Konference

Interspeech Conference

Vydavatel

International Speech Communication Association

Místo

Graz

DOI

10.21437/Interspeech.2019-2842

UT WoS

000831796403001

EID Scopus

2-s2.0-85074683253

BibTeX

@inproceedings{BUT159999,
  author="STAFYLAKIS, T. and ROHDIN, J. and PLCHOT, O. and MIZERA, P. and BURGET, L.",
  title="Self-supervised speaker embeddings",
  booktitle="Proceedings of Interspeech",
  year="2019",
  journal="Proceedings of Interspeech",
  volume="2019",
  number="9",
  pages="2863--2867",
  publisher="International Speech Communication Association",
  address="Graz",
  doi="10.21437/Interspeech.2019-2842",
  issn="1990-9772",
  url="https://www.isca-speech.org/archive/Interspeech_2019/pdfs/2842.pdf"
}

Soubory

pdf stafylakis_is2019_192842.pdf 320 kB

Projekty

Dolování infoRmAcí z řeči Pořízené vzdÁlenými miKrofony, MV, Bezpečnostní výzkum České republiky 2015-2020, VI20152020025, zahájení: 2015-10-01, ukončení: 2020-09-30, ukončen
IT4Innovations excellence in science, MŠMT, Národní program udržitelnosti II, LQ1602, zahájení: 2016-01-01, ukončení: 2020-12-31, ukončen
Neuronové reprezentace v multimodálním a mnohojazyčném modelování, GAČR, Grantové projekty exelence v základním výzkumu EXPRO - 2019, GX19-26934X, zahájení: 2019-01-01, ukončení: 2023-12-31, ukončen
Neuronové sítě shrnující sekvence pro rozpoznávání mluvčího, EU, Horizon 2020, 5SA15094, zahájení: 2016-07-01, ukončení: 2019-06-30, ukončen
Zpracování, zobrazování a analýza multimediálních a 3D dat, VUT, Vnitřní projekty VUT, FIT-S-17-3984, zahájení: 2017-03-01, ukončení: 2020-02-29, ukončen

Výzkumné skupiny

Výzkumná skupina dolování dat z řeči BUT Speech@FIT (VZ SPEECH)

Pracoviště

Ústav počítačové grafiky a multimédií (UPGM)