Comparative Performance of Large Language Models: ChatGPT o۳-mini and Deepseek R۱ in Pediatric Hematology/Oncology Questions

سال انتشار: 1404
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 210

متن کامل این مقاله منتشر نشده است و فقط به صورت چکیده یا چکیده مبسوط در پایگاه موجود می باشد.
توضیح: معمولا کلیه مقالاتی که کمتر از ۵ صفحه باشند در پایگاه سیویلیکا اصل مقاله (فول تکست) محسوب نمی شوند و فقط کاربران عضو بدون کسر اعتبار می توانند فایل آنها را دریافت نمایند.

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

AIMS02_320

تاریخ نمایه سازی: 29 تیر 1404

چکیده مقاله:

Background and Aims: This study compares the performance of two large language models, ChatGPT o۳-mini and Deepseek R۱, in answering pediatric hematology/oncology multiple-choice questions. With the growing integration of artificial intelligence in clinical decision-making, it is essential to assess the ability of these models to accurately process complex medical queries. The study aims to evaluate their accuracy, response time, and reasoning capabilities under standardized conditions. Methods: A cross-sectional analysis was performed using ۱۰۰ self-assessment multiple-choice questions originally developed by the American Society of Pediatric Hematology/Oncology. Due to the inclusion of images in some items, eight questions were excluded, resulting in ۹۲ paired items for evaluation. Both models were tested under controlled conditions with memory features disabled to ensure independent processing of each question. Performance metrics included the accuracy of response selection, response time measured from query submission to answer output, and qualitative assessment of the reasoning process. IBM SPSS version ۲۷ was utilized for the statistical analyses. Statistical analyses, including McNemar’s test and the Mann-Whitney U test, were applied to identify significant differences between the models. Results: Deepseek R۱ demonstrated a significantly higher accuracy rate (approximately ۸۵.۸۷%) compared to ChatGPT o۳-mini (۶۵.۵%). Although ChatGPT o۳-mini provided faster response times, its performance in processing complex clinical scenarios was less consistent. The statistical analyses confirmed that the differences in both accuracy and response times between the models were significant (p < ۰.۰۰۱). Conclusion: The findings indicate that while both models show promise for supporting clinical decision-making in pediatric hematology/oncology, Deepseek R۱ offers superior accuracy and more reliable clinical reasoning, despite its slower response time. These results suggest that further research is warranted to optimize the trade-off between speed and precision, and to evaluate the applicability of these models in real-world clinical settings.

نویسندگان

Sahel Sharifpoor Saleh

Master of Science Student in Medical Informatics, Faculty of Paramedical Sciences, Urmia University of Medical Sciences, Urmia, Iran

Sara Mohammadi Kalashani

Bachelor of Science in Health Information Technology, School of Paramedical Sciences, Urmia University of Medical Sciences, Urmia, Iran

Mohammad Zarbi

Master of Medical Informatics, Management of Statistics and Information Technology, Urmia University of Medical Sciences, Urmia, Iran