inklap

Adaptive Contrastive Cross-Modal Transformer for Enhanced Text-to-Image Person Re-Identification

, Tsoy Yelizaveta, Rahman Touhidur, , Bishwas Shankar Palikhe, · International Journal of Innovative Research in Computer and Communication Engineering · 2025

Traditional person re-identification systems relying solely on visual features struggle with real-world challenges including occlusions, illumination variations, viewpoint changes, and appearance modifications, resulting in poor generalization across diverse surveillance scenarios. This paper presents the Adaptive Contrastive Cross-Modal Transformer (AC-CMT), a novel multimodal framework that synergistically combines visual and textual modalities for robust person re-identification. The AC-CMT architecture combines a Vision Transformer (ViT) backbone for finegrained visual feature extraction, a CLIP-based text encoder for semantic understanding, and an innovative adaptive cross-modal attention mechanism that dynamically aligns image regions with corresponding textual descriptions. The model incorporates two new loss functions: Similarity-Based Alignment Loss (SBAL) for semantic feature alignment and Temperature-Scaled Cross-Modal Matching with Dynamic Weights (TSCM-DW) for handling challenging alignment situations. Comprehensive experiments on CUHK-PEDES and ICFG-PEDES benchmark datasets show significant performance improvements, achieving 71.44% and 61.51% Rank-1 accuracy respecti

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً