inklap

EGRTE: adversarially training a self-explaining smoothed classifier for certified robustness

Zijin Lin, Jinwen He, Yue Zhao, Ruigang Liang, Hu Li, Zhendong Wu · Cybersecurity · 2025

Abstract Deep learning has transformed fields such as computer vision, natural language processing, and audio analysis through its powerful pattern recognition and predictive capabilities. However, the robustness of these models remains a major concern, as they are highly vulnerable to adversarial attacks-subtle, intentional perturbations that lead to incorrect predictions. While recent defenses like adversarial training and defensive distillation aim to improve robustness, they have notable drawbacks, including overfitting and degraded performance under strong attacks. Certified defenses, such as robust training and Randomized Smoothing, offer theoretical guarantees within a specific perturbation radius, yet struggle to reflect real-world robustness due to efficiency bottlenecks and the unpredictable nature of actual adversarial attacks. These challenges reveal a critical gap between current defenses and real-world attack scenarios, highlighting the need for more practical and resilient solutions. To address the challenges of defense-attack gaps and the inefficiency in robust training, we introduce the Explanation-Guided Robust Training Enhancer (EGRTE). EGRTE combines a

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً