Result Details
Bayesian Neural Representation Augmentations for Domain Generalization in Speaker Verification and Anti-Spoofing
Mak Man-Wai
Rohdin Johan Andréas, M.Sc., Ph.D., FIT (FIT), DCGM (FIT)
Plchot Oldřich, Ing., Ph.D., DCGM (FIT)
Lee Kong Aik
Heřmanský Hynek, prof. Ing., Dr. Eng., DCGM (FIT)
Speaker verification and anti-spoofing systems often face significant performance degradation when encountering out-of-domain (OOD) scenarios due to the variability in speaker characteristics, recording environments, and unknown spoofing attacks. Domain generalization (DG) methods are designed to identify features that remain consistent across various domains, thereby enhancing the robustness of deep learning models. Style augmentation is a DG approach that synthesizes novel domain features from feature statistics. However, previous works only consider the covariances within individual mini-batches during style augmentation, which may lead to outliers in the augmented samples. In this paper, we propose to improve the generalization capability of speaker verification and anti-spoofing systems by leveraging the posterior covariances of neural representations in the outputs of Transformer layers of a neural network through Bayesian adaptation and call the method Bayesian Neural Representation Augmentation (BNRA). We use the covariances estimated from the current mini-batch to define the covariances of a Gaussian likelihood function for the observed data. The covariances of the prior distribution of latent factors are updated recursively. Our plug-and-play module can be seamlessly integrated into existing self-supervised learning networks without extra parameters. It is the first to incorporate augmentation of neural representations for OOD generalization in speaker verification and anti-spoofing, with extensive experiments confirming its consistent performance gains.
Diodes, Light emitting diodes, Superluminescent diodes, Protocols, Deepfakes, Videos, Communication systems, Codecs, Telephony, DSL, Domain-invariant speaker verification, antispoofing, domain generalization, Bayesian learning
@article{BUT212047,
author="Jin {Li} and {} and Johan Andréas {Rohdin} and Oldřich {Plchot} and {} and Hynek {Heřmanský}",
title="Bayesian Neural Representation Augmentations for Domain Generalization in Speaker Verification and Anti-Spoofing",
journal="IEEE transactions on audio, speech, and language processing",
year="2026",
volume="34",
number="2026",
pages="2198--2212",
doi="10.1109/TASLPRO.2026.3682044",
url="https://ieeexplore.ieee.org/document/11477098"
}