inklap

Evaluation of Reliability, Repeatability, Robustness, and Confidence of GPT-3.5 and GPT-4 on a Radiology Board–style Examination

Satheesh Krishna, Nishaant Bhambra, Robert Bleakney, Rajesh Bhayana · Radiology · 2024

Default GPT-3.5 and GPT-4 were reliably accurate over three attempts at answering radiology board–style multiple-choice questions but had poor repeatability and robustness and were frequently overconfident, limiting usability without domain-specific optimization.

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً