inklap

Towards optimal adversarial texts: character, word, and sentence

Pengchuan Wang, Deqiang Li, Qianmu Li · Cybersecurity · 2025

Abstract Natural language processing models are widely acknowledged for their strong data fitting capabilities, diverse application scenarios, and adaptable learning methodologies. However, these models, including the large language models, exhibit sensitivity to adversarial example attacks. These examples are slightly perturbed from the pristine text but mislead the model classification. Nevertheless, the existing attack methods primarily focus on the attack effectiveness without semantics-preservation considered. Moreover, the trade-off between evasion effectiveness and concealment of perturbed texts is less investigated. In this study, we propose a multi-objective adversarial text generation framework (MOATG) that simultaneously optimizes attack success rate, imperceptibility, and semantic similarity. Tailored objective functions and dominance relations are designed for character-, word-, and sentence-level perturbations. MOATG is evaluated against five baselines across five benchmark datasets. Experimental results show that MOATG achieves a 5.38% average improvement in attack success rate and reduces word error rate by 1.48%, demonstrating its effectiveness in

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً