Automated crime classification is critical for law enforcement resource allocation, yet crime datasets exhibit severe class imbalance and fairness issues rooted in historical policing patterns. This paper presents a methodologically rigorous machine learning framework addressing these challenges through three strategies: adaptive class weighting, SMOTE-NC (for mixed categorical/numerical tabular data), and SMOTE-NC with Tomek link removal. Models are evaluated via nested spatiotemporal cross-validation-spatial block outer loop combined with forward-chaining inner loop-on three independent datasets: the Rajshahi Metropolitan Police (RMP) incident dataset, San Francisco, and Chicago. We compare tabular-native architectures (XGBoost, CatBoost, TabNet) with in-processing fairness baselines (FairXGBoost, FairGBM). Protected demographic attributes are strictly excluded from training features and reserved solely for post-hoc fairness auditing. FairXGBoost with SMOTE-NC achieved the best fairness-accuracy tradeoff (92.4% accuracy, 0.901 macro F1, p < 0.001) with an 8.1% minority class recall improvement and a demographic parity gap reduced to 5.8%. We explicitly discuss the Impossibilit
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً