Result Details

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

UDUPA, S.; WATANABE, S.; SCHWARZ, P.; CERNOCKY, J. Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training. In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). Honolulu, Hawaii Islands, USA: IEEE, 2025. p. 1-8. ISBN: 979-8-3315-4426-3.
Type
conference paper
Language
English
Authors
Udupa Sathvik, Ing., FIT (FIT), DCGM (FIT)
Watanabe Shinji
Schwarz Petr, Ing., Ph.D., DCGM (FIT)
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Abstract

Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.

Keywords

endpointing; turn-taking prediction

URL
Published
2025
Pages
1–8
Proceedings
2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Conference
2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
ISBN
979-8-3315-4426-3
Publisher
IEEE
Place
Honolulu, Hawaii Islands, USA
DOI
EID Scopus
BibTeX
@inproceedings{BUT212025,
  author="Sathvik {Udupa} and  {} and Petr {Schwarz} and Jan {Černocký}",
  title="Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training",
  booktitle="2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)",
  year="2025",
  pages="1--8",
  publisher="IEEE",
  address="Honolulu, Hawaii Islands, USA",
  doi="10.1109/asru65441.2025.11434752",
  isbn="979-8-3315-4426-3",
  url="https://ieeexplore.ieee.org/document/11434752"
}
Files
Projects
Research groups
Departments
Back to top