Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study.
L’essentiel
Timely and highly accurate diagnoses by physicians play a crucial role in improving the quality and effectiveness of patient treatment outcomes. Currently, the use of artificial intelligence capabilities in this area has garnered the attention of many health science researchers. Therefore, the main goal of this preliminary study was to compare the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in definitive and differential diagnoses using standardized clinical vignettes. This descriptive comparative study evaluated the diagnostic accuracy of 10 emergency medicine physicians and 4 large language models (LLMs)-ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1-using 10 standardized clinical vignettes. All LLMs were accessed via official web interfaces. Clinical vignettes were developed from emergency department presentations, validated by an expert panel, and presented as text-only inputs to all evaluators. Diagnostic accuracy was assessed using a standardized scoring protocol: definitive diagnoses required an exact match with expert-derived reference standards, while differential diagnoses required ≥ 3 matches. Overall diagnostic accuracy was the primary outcome. Data were analyzed using Pearson's Chi-square test, McNemar's test, and Generalized Estimating Equation (GEE) logistic regression with Bonferroni correction (SPSS version 28). A total of 280 diagnostic evaluations (10 clinical cases assessed by 14 evaluators across 2 diagnosis types) were analyzed. Overall diagnostic accuracy was 61.79%. AI models demonstrated significantly higher overall accuracy (73.75%) compared to emergency medicine physicians (57.00%, p = 0.014). Across all evaluators, definitive diagnoses were more accurate than differential diagnoses (70.71% vs. 52.86%). Generalized Estimating Equation (GEE) analysis revealed a significant interaction between evaluator group and diagnosis type (p = 0.036). Specifically, physicians experienced a significant decline in accuracy when providing differential diagnoses compared to definitive diagnoses (45.0% vs. 69.0%, p = 0.003; remained significant after Bonferroni correction). In contrast, AI models maintained consistently high accuracy across both diagnosis types, with no significant difference between definitive (75.0%) and differential (72.5%) diagnoses (p = 1.000). Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks. While human physicians struggled significantly with differential diagnoses, AI models maintained high and stable performance regardless of the diagnostic complexity. These findings indicate that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool, particularly in complex scenarios requiring differential diagnostic reasoning. Due to study limitations, such as the small number of clinical scenarios and assessors, these findings should be interpreted with significant caution.
Synthèse détaillée
Résumé original
Timely and highly accurate diagnoses by physicians play a crucial role in improving the quality and effectiveness of patient treatment outcomes. Currently, the use of artificial intelligence capabilities in this area has garnered the attention of many health science researchers. Therefore, the main goal of this preliminary study was to compare the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in definitive and differential diagnoses using standardized clinical vignettes. This descriptive comparative study evaluated the diagnostic accuracy of 10 emergency medicine physicians and 4 large language models (LLMs)-ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1-using 10 standardized clinical vignettes. All LLMs were accessed via official web interfaces. Clinical vignettes were developed from emergency department presentations, validated by an expert panel, and presented as text-only inputs to all evaluators. Diagnostic accuracy was assessed using a standardized scoring protocol: definitive diagnoses required an exact match with expert-derived reference standards, while differential diagnoses required ≥ 3 matches. Overall diagnostic accuracy was the primary outcome. Data were analyzed using Pearson's Chi-square test, McNemar's test, and Generalized Estimating Equation (GEE) logistic regression with Bonferroni correction (SPSS version 28). A total of 280 diagnostic evaluations (10 clinical cases assessed by 14 evaluators across 2 diagnosis types) were analyzed. Overall diagnostic accuracy was 61.79%. AI models demonstrated significantly higher overall accuracy (73.75%) compared to emergency medicine physicians (57.00%, p = 0.014). Across all evaluators, definitive diagnoses were more accurate than differential diagnoses (70.71% vs. 52.86%). Generalized Estimating Equation (GEE) analysis revealed a significant interaction between evaluator group and diagnosis type (p = 0.036). Specifically, physicians experienced a significant decline in accuracy when providing differential diagnoses compared to definitive diagnoses (45.0% vs. 69.0%, p = 0.003; remained significant after Bonferroni correction). In contrast, AI models maintained consistently high accuracy across both diagnosis types, with no significant difference between definitive (75.0%) and differential (72.5%) diagnoses (p = 1.000). Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks. While human physicians struggled significantly with differential diagnoses, AI models maintained high and stable performance regardless of the diagnostic complexity. These findings indicate that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool, particularly in complex scenarios requiring differential diagnostic reasoning. Due to study limitations, such as the small number of clinical scenarios and assessors, these findings should be interpreted with significant caution.