Result Details
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
Kesiraju Santosh, Ph.D., DCGM (FIT)
Beneš Karel
Yusuf Bolaji, DCGM (FIT)
Burget Lukáš, doc. Ing., Ph.D., DCGM (FIT)
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
This paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and generalization in both in- and out-of-domain settings. The proposed method, Decoder-Centric Regularization in EncoderDecoder (DeCRED), adds auxiliary classifiers to the decoder, enabling next token prediction via intermediate logits. Empirically, DeCRED reduces the mean internal LM BPE perplexity by 36.6% relative to 11 test sets. Furthermore, this translates into actual WER improvements over the baseline in 5 of 7 in-domain and 3 of 4 out-of-domain test sets, reducing macro WER from 6.4% to 6.3% and 18.2 % to 16.2 %, respectively. On TEDLIUM3, DeCRED achieves 7.0 % WER, surpassing the baseline and encoder-centric InterCTC regularization by 0.6 % and 0.5%, respectively. Finally, we compare DeCRED with OWSM v3.1 and Whisper-medium, showing competitive WERs despite training on much less data with fewer parameters.
intermediate regularization | out-of-domain generalization | speech recognition
@inproceedings{BUT212028,
author="Alexander {Polok} and Santosh {Kesiraju} and Karel {Beneš} and Bolaji {Yusuf} and Lukáš {Burget} and Jan {Černocký}",
title="DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition",
booktitle="ASRU 2025 2025 IEEE Automatic Speech Recognition and Understanding Workshop",
year="2025",
pages="7",
publisher="Institute of Electrical and Electronics Engineers Inc.",
address="Honolulu, Hawaii Islands, USA",
doi="10.1109/ASRU65441.2025.11434661",
isbn="979-8-3315-4426-3",
url="https://ieeexplore.ieee.org/document/11434661"
}