inklap

InfoShadow: NTK & MI guided adversarial attacks on speaker identification systems

Ruixin Song, Youliang Tian, Mengqian Li, Ze Yang, Ruohan Wang · Cybersecurity · 2025

Abstract Adversarial attacks on speaker identification (SI) systems have become a critical security concern, particularly in targeted black-box scenarios where access to the target model is limited. This paper proposes a novel framework that creates highly transferable adversarial examples. We use a voice conversion (VC) model to synthesize shadow data from a single target speech sample, which is then used to train two diverse surrogate models. Neural Tangent Kernel (NTK) theory is employed to align acoustic feature spaces, while mutual information optimization enforces consistency between the surrogate models’ predictions. Consequently, the adversarial attack is formulated as a min-max game that maximizes attack success while preserving speech quality. Extensive experiments on LibriSpeech and VCTK datasets demonstrate that our method significantly improves the transferability and effectiveness of adversarial examples compared to conventional approaches. Our findings suggest that generating shadow data through voice conversion followed by surrogate model training under information-theoretic constraints is a promising strategy for robust adversarial attacks.

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً