Robotic waste sorting presents significant challenges, including object variability, cluttered environments, and the predominant reliance on deep learning and traditional computer vision techniques, which typically demand extensive datasets and task-specific training. This paper introduces a robotic waste sorting system that integrates the Gemini Vision–Language–Action (VLA) model with a KUKA LBR iiwa collaborative robot and an RGB-D camera. Our approach leverages the advanced reasoning capabilities of large, pre-trained VLA models to perform waste sorting, without requiring explicit training or dataset collection. Key contributions include the development of effective prompt engineering strategies for waste object identification, the assessment of the VLA’s performance in terms of inference time and accuracy, and the development of different grasping strategies for operation in cluttered scenarios. Our experimental tests demonstrated that the system’s inference time is between 2 and 4 s, which is suitable for collaborative robotic applications, and the system achieved a high overall classification accuracy of 89.64%. Crucially, we demonstrated that integration of RGB-D sensing enh
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً