Detail výsledku

BCN2BRNO: ASR System Fusion for Albayzin 2020 Speech to Text Challenge

KOCOUR, M.; CÁMBARA, G.; LUQUE, J.; BONET, D.; FARRÚS, M.; KARAFIÁT, M.; VESELÝ, K.; ČERNOCKÝ, J. BCN2BRNO: ASR System Fusion for Albayzin 2020 Speech to Text Challenge. Proceedings of IberSPEECH 2021. Vallaloid: International Speech Communication Association, 2021. p. 113-117.
Typ
článek ve sborníku konference
Jazyk
anglicky
Autoři
Kocour Martin, Ing., UPGM (FIT)
CÁMBARA, G.
Luque Jordi, FIT (FIT)
BONET, D.
FARRÚS, M.
Karafiát Martin, Ing., Ph.D., UPGM (FIT)
Veselý Karel, Ing., Ph.D., UPGM (FIT)
Černocký Jan, prof. Dr. Ing., UPGM (FIT)
Abstrakt

This paper describes the joint effort of BUT and Telefónica Researchon the development of Automatic Speech Recognitionsystems for the Albayzin 2020 Challenge. We compare approachesbased on either hybrid or end-to-end models. In hybridmodelling, we explore the impact of a SpecAugment layeron performance. For end-to-end modelling, we used a convolutionalneural network with gated linear units (GLUs). Theperformance of such model is also evaluated with an additionaln-gram language model to improve word error rates. We furtherinspect source separation methods to extract speech fromnoisy environments (i.e. TV shows). More precisely, we assessthe effect of using a neural-based music separator named Demucs.A fusion of our best systems achieved 23.33% WER inofficial Albayzin 2020 evaluations. Aside from techniques usedin our final submitted systems, we also describe our efforts inretrieving high-quality transcripts for training.

Klíčová slova

fusion, end-to-end model, hybrid model, semisupervised,automatic speech recognition, convolutional neuralnetwork.

URL
Rok
2021
Strany
113–117
Sborník
Proceedings of IberSPEECH 2021
Konference
IberSPEECH 2021 Conference
Vydavatel
International Speech Communication Association
Místo
Vallaloid
DOI
BibTeX
@inproceedings{BUT175823,
  author="KOCOUR, M. and CÁMBARA, G. and LUQUE, J. and BONET, D. and FARRÚS, M. and KARAFIÁT, M. and VESELÝ, K. and ČERNOCKÝ, J.",
  title="BCN2BRNO: ASR System Fusion for Albayzin 2020 Speech to Text Challenge",
  booktitle="Proceedings of IberSPEECH 2021",
  year="2021",
  pages="113--117",
  publisher="International Speech Communication Association",
  address="Vallaloid",
  doi="10.21437/IberSPEECH.2021-24",
  url="https://www.isca-speech.org/archive/iberspeech_2021/kocour21_iberspeech.html"
}
Soubory
Projekty
Automatický sběr a zpracování hlasových dat z letecké komunikace, EU, Horizon 2020, zahájení: 2019-11-01, ukončení: 2022-02-28, ukončen
Multi-lingualita v řečových technologiích, MŠMT, INTER-EXCELLENCE - Podprogram INTER-ACTION, LTAIN19087, zahájení: 2020-01-01, ukončení: 2023-08-31, ukončen
Neuronové reprezentace v multimodálním a mnohojazyčném modelování, GAČR, Grantové projekty exelence v základním výzkumu EXPRO - 2019, GX19-26934X, zahájení: 2019-01-01, ukončení: 2023-12-31, ukončen
Výzkumné skupiny
Pracoviště
Nahoru