Abstract The growing deployment of vehicles equipped with driver assistance and automated driving systems presents new challenges for crash record classification, as existing police-reported databases vary considerably in the completeness and consistency of automation-level metadata. This study benchmarks five tabular machine learning and deep learning models: random forest, XGBoost, MambaAttention, prior-data fitted network (TabPFN), and TabTransformer for classifying reported SAE automation categories using structured crash records from the Texas CRIS database (2024), comprising 4649 records across assisted driving (SAE Level 1), partial automation (SAE Level 2), and advanced automation (SAE Levels 3–5). The SAE automation label in each record reflects the vehicle’s designed automation capability derived from make, model, and year specifications rather than confirmed system engagement at crash time; the classification task, therefore, addresses SAE-coded vehicle capability identification rather than active automation detection. Records were partitioned at the crash level using GroupShuffleSplit to prevent crash-level data leakage, yielding 3
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً