Faculty of Information Technology, BUT

Publication Details

How To Improve Your Speaker Embeddings Extractor in Generic Toolkits

ZEINALI Hossein, BURGET Lukáš, ROHDIN Johan A., STAFYLAKIS Themos and ČERNOCKÝ Jan. How To Improve Your Speaker Embeddings Extractor in Generic Toolkits. In: Proceedings of ICASSP 2019. Brighton: IEEE Signal Processing Society, 2019, pp. 6141-6145. ISBN 978-1-5386-4658-8. Available from: https://ieeexplore.ieee.org/abstract/document/8683445
Czech title
Jak zlepšit Váš extraktor embeddingů mluvčích v běžných toolkitech
Type
conference paper
Language
english
Authors
Zeinali Hossein, Ph.D. (DCGM FIT BUT)
Burget Lukáš, doc. Ing., Ph.D. (DCGM FIT BUT)
Rohdin Johan A., Dr. (DCGM FIT BUT)
Stafylakis Themos (OMILIA)
Černocký Jan, doc. Dr. Ing. (DCGM FIT BUT)
URL
Keywords
Deep neural network, speaker embedding, xvector, Tensorflow, Kaldi.
Abstract
Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine several tricks in training, such as the effects of normalizing input features and pooled statistics, different methods for preventing overfitting as well as alternative nonlinearities that can be used instead of Rectifier Linear Units. In addition, we investigate the difference in performance between TDNN and CNN, and between two types of attention mechanism. Experimental results on Speaker in the Wild, SRE 2016 and SRE 2018 datasets demonstrate the effectiveness of the proposed implementation.
Published
2019
Pages
6141-6145
Proceedings
Proceedings of ICASSP 2019
Conference
International Conference on Acoustics, Speech, and Signal Processing, Brighton, GB
ISBN
978-1-5386-4658-8
Publisher
IEEE Signal Processing Society
Place
Brighton, GB
BibTeX
@INPROCEEDINGS{FITPUB12037,
   author = "Hossein Zeinali and Luk\'{a}\v{s} Burget and A. Johan Rohdin and Themos Stafylakis and Jan \v{C}ernock\'{y}",
   title = "How To Improve Your Speaker Embeddings Extractor in Generic Toolkits",
   pages = "6141--6145",
   booktitle = "Proceedings of ICASSP 2019",
   year = 2019,
   location = "Brighton, GB",
   publisher = "IEEE Signal Processing Society",
   ISBN = "978-1-5386-4658-8",
   language = "english",
   url = "https://www.fit.vut.cz/research/publication/12037"
}
Back to top