Default GPT-3.5 and GPT-4 were reliably accurate over three attempts at answering radiology board–style multiple-choice questions but had poor repeatability and robustness and were frequently overconfident, limiting usability without domain-specific optimization.
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً