inklap

Beyond technical adequacy: A holistic framework for evaluating artificial intelligence systems through the scaling responsible artificial intelligence mentorship approach

Caroline Gans Combe · Evaluation · 2025

Artificial intelligence evaluation practices face a fundamental challenge: traditional technical measures cannot adequately capture the complex socio-organizational impacts of artificial intelligence systems in real-world contexts. This research bridges the critical gap between technical evaluation and holistic evaluation by exploring how insights from evaluation theory can inform artificial intelligence system assessment. Using a mixed-methods case approach inspired by developmental evaluation, realistic evaluation, and empowerment evaluation principles, we analyzed 12 artificial intelligence implementations in organizations within the scaling responsible artificial intelligence mentoring program. Data collection consisted of semi-structured participant interviews, participant observation, and data extraction from various institutional databases, all guided by specific evaluation theoretical frameworks. Our analysis reveals five generative principles for a holistic evaluation of artificial intelligence: epistemic pluralism, democratic authority, contextual responsiveness, temporal sensitivity, and reflexive critique. These principles transcend procedural criteria to encompass fund

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً