Collaborative perception can mitigate occlusion and range limitations in autonomous driving, but deployment remains constrained by strict bandwidth budgets and heterogeneous agent stacks. We propose a communication-efficient and backbone-agnostic framework in which each agent’s encoder is treated as a black box, and a lightweight interpreter maps its intermediate features into a canonical space. To reduce transmission cost, we integrate codebook-based compression that sends only compact discrete indices, while a prompt-guided decoder reconstructs semantically aligned features on the ego vehicle for downstream fusion. Training follows a two-phase strategy: Phase 1 jointly optimizes interpreters, prompts, and fusion components for a fixed set of agents; Phase 2 enables plug-and-play onboarding of new agents by tuning only their specific prompts. Experiments on OPV2V and OPV2VH+ show that our method consistently outperformed early-, intermediate-, and late-fusion baselines under equal or lower communication budgets. With a codebook of size 128, the proposed pipeline preserved over 95% of the uncompressed detection accuracy while reducing communication cost by more than two orders of m
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً