Result Details
Trainable Multi-Channel Front-Ends for Joint Beamforming and Speaker Embedding Extraction
Plchot Oldřich, Ing., Ph.D., DCGM (FIT)
Burget Lukáš, doc. Ing., Ph.D., DCGM (FIT)
Zhang Chunlei
Černocký Jan, prof. Dr. Ing., DCGM (FIT)
Yu Meng
Multi-channel speaker verification (SV), employing numerous microphones for capturing enrollment and/or test recordings, gained attention for its benefits in far-field scenarios. While some studies approach the problem by designing multi-channel embedding extractors, we focus on building and thoroughly analyzing a framework integrating beamforming pre-processing paired with single-channel embedding extraction. This strategy benefits from accommodating both multi-channel and single-channel inputs. Furthermore, it provides human-interpretable intermediate output — enhanced speech —
that can be independently evaluated and related to SV performance. We first focus on the front-end, taking advantage of deep-learning source separation for direct or indirect mask estimation required by the beamformer. We alternate single-channel network architectures, subsequently extended to multi-channel ones by reference channel attention (RCA). We also analyze the impact of beamformer and network output fusion. Finally, we show improvements brought by end-to-end fine-tuning the entire architecture facilitated by our newly designed multi-channel corpus, MultiSV2, extending our previous MultiSV dataset.
Multi-channel speaker verification, beamforming, MultiSV, reference channel attention
@article{BUT200198,
author="Ladislav {Mošner} and Oldřich {Plchot} and Lukáš {Burget} and {} and Jan {Černocký} and {}",
title="Trainable Multi-Channel Front-Ends for Joint Beamforming and Speaker Embedding Extraction",
journal="COMPUTER SPEECH AND LANGUAGE",
year="2026",
volume="99",
number="101944",
pages="1--46",
doi="10.1016/j.csl.2026.101944",
issn="0885-2308",
url="https://www.sciencedirect.com/science/article/pii/S0885230826000070"
}