DÄ internationalArchive4/2026Evaluation of Online Large Language Models as Decision Support for the Management of Abdominal Pain in the Emergency Department

Research letter

Evaluation of Online Large Language Models as Decision Support for the Management of Abdominal Pain in the Emergency Department

A Pilot Study

Dtsch Arztebl Int 2026; 123: 117-8. DOI: 10.3238/arztebl.m2025.0237

Henn, J; Feodorovici, P; Dohmen, J; Kalff, J C; Gräff, I; Matthaei, H

LNSLNS

Abdominal pain is a common yet challenging presenting symptom in emergency medicine due to its heterogeneous causes and comparatively high mortality rate (1).

The initial suspected diagnosis and its urgency play a key role in the rapid and targeted evaluation (2).

Artificial intelligence (AI), and large language models (LLMs) in particular, are under consideration as decision-support tools. No corresponding analysis in abdominal pain has been conducted to date.

Methods

Based on real, manually anonymized case vignettes from University Hospital Bonn, Germany, ChatGPT Plus (GPT-4o), Gemini (1.5 Pro), specialist visceral surgeons, and final-year medical students were compared with
regard to diagnostic accuracy and urgency assessment.

The actual diagnosis and need for surgery within 24 h served as the reference.

Evaluators were presented with batches of 10 randomly selected cases, each of which was fully assessed. All results are presented as median with interquartile range.

Results

A total of 100 case vignettes were created from 52 female and 48 male patients (median age 49 years [29.8–62.3]). The chatbots assessed 50 cases in 10 iterations each, meaning that each case was evaluated a median of five times. Eight specialists from three centers evaluated a median of 35 cases (20–50) and 14 medical students from six centers a median of 40 cases (10–50). Each case was assessed a median of three times by specialists, 4.5 times by medical students, and at least once by each group. In terms of accuracy (top-1 diagnosis), ChatGPT (0.44; [0.41–0.49]), Gemini (0.43; [0.41–0.48]), and specialists (0.44; [0.43–0.46]) achieved similar median values, whereas medical students performed worse (0.39; [0.31–0.49]). ChatGPT gave the reference diagnosis (median 0.76; [0.73–0.78]) more often under the top-3 diagnoses compared to Gemini (0.66; [0.65–0.70]), the specialists (0.66; [0.61–0.68]), and the medical students (0.61; [0.51–0.70]).

Compared to the reference (51 urgent and 49 non-urgent cases), all evaluators classified more cases as urgent (median: ChatGPT 72; [71–74], Gemini 74; [71–76], specialists 63; [56–71], medical students 69; [64–80]). While ChatGPT (0.79; [0.77–0.82]), Gemini (0.78; [0.74–0.79]), and the specialists (0.75; [0.72–0.80]) achieved similar F1 scores, the medical students performed less precisely (0.68; [0.67–0.72]). Here, ChatGPT demonstrated higher sensitivity (1.00; [0.97–1.00]) compared to Gemini (0.94; [0.90–0.96]), specialists (0.86; [0.77–0.92]), and medical students (0.80; [0.72–0.93]). In contrast, specialists achieved the highest median specificity (0.64; [0.48–0.68]), ahead of ChatGPT (0.53; [0.47–0.58]), Gemini (0.47; [0.43–0.50]), and medical students (0.44; [0.33–0.50]) (Figure). Both chatbots assessed the cases a median of approximately 10 times faster than specialists and nearly 18 times faster than students.

Direct comparison of sensitivity and specificity across evaluator groups.
Figure
Direct comparison of sensitivity and specificity across evaluator groups.

Discussion

All groups suspected specific diagnoses more frequently than indicated by the reference. This approach appears beneficial for clinical decision-making, given that a specific suspected diagnosis is needed for further radiological evaluation (urgency and modality) (2). None of the groups achieved a diagnostic accuracy (top-1) of over 45%, underscoring the complexity of this presenting symptom. As expected, medical students were less accurate, while ChatGPT performed at a comparable level to specialists, and even outperformed them in the top-3 differential diagnoses.

On the one hand, our observations are in line with retrospective single-center data from Munich, Germany, in which GPT-4 outperformed resident physicians in the assessment of internal medicine emergencies (3). On the other hand, the diagnostic accuracy of a proprietary AI application in the prospective, double-blind eRadaR trial was inferior to that of the classic patient–physician interaction (4). The differences can be explained by heterogeneous methodological approaches, while standardized prospective comparisons are lacking.

Prompt assessment and treatment of urgent diagnoses can significantly reduce complications (4). In the assessment of urgency, ChatGPT outperformed all comparison groups with a sensitivity of 100%. How the model resolves the trade-off between sensitivity and specificity remains unclear, and its proprietary nature precludes further analysis. In view of demographic change and ongoing structural reforms, it is becoming increasingly unrealistic for every patient to be primarily assessed by a specialist physician (5).

Final-year medical students are preparing to provide emergency care as resident physicians. Our results suggest that they may benefit from LLM-based decision support by comparing their own assessments with those of the LLM. Although this approach could conceivably ensure a minimum standard in terms of quality assurance, the clinical consequences and overall benefit remain unclear and require dedicated scientific evaluation. Ethical implications also need to be considered. For example, the transparency of online LLMs is limited, and outsourcing differential diagnostic assessments may create dependence on LLM decision support. Remarkably, ChatGPT achieved this performance without examples (zero-shot). Although the single-center and retrospective design of our case vignettes limits their generalizability, the study involved real cases that were not part of the LLMs’ training data, and data leakage can be ruled out. Since LLMs are trained on general data, they are less susceptible to overfitting than specially trained AI applications.

Our data suggest that LLMs can be used as decision support in the emergency department. Looking ahead, fine-tuning with local data may be beneficial for improving performance, control, and data protection. Against the backdrop of increasing gaps in the provision of care, further scientific investigation in clinical studies appears warranted.

Jonas Henn, Philipp Feodorovici, Jonas Dohmen, Jörg C. Kalff, Ingo Gräff, Hanno Matthaei

Ethics statement

The study was approved by the Ethics Committee of University Hospital Bonn (163/23-EP), and explicit informed consent was not required.

Acknowledgments

We would like to thank all participants for their time and effort in assessing the case vignettes.

Funding

This study was funded in part by the Ministry of Economic Affairs, Industry, Climate Action and Energy of the State of North Rhine-Westphalia (Ministerium für Wirtschaft, Industrie, Klimaschutz und Energie des Landes Nordrhein-Westfalen; project: Innovative Secure Medical Campus).

Conflict of interest statement

The remaining authors declare that no conflict of interest exists.

Manuscript submitted on 10 July 2025, revised version accepted on
15 December 2025.

Translated from the original German by Christine Rye.

Cite this as
Henn J, Feodorovici P, Dohmen J, Kalff JC, Gräff I, Matthaei H: Evaluation of online large language models as decision support for the management of abdominal pain in the emergency department: A pilot study. Dtsch Arztebl Int 2026; 123: 117–8. DOI: 10.3238/arztebl.m2025.0237

1.
Helbig L, Möckel M, Fischer-Rosinsky A, Slagman A: Non-traumatic ­abdominal pain—a retrospective analysis of secondary data from 448 689 cases treated in two emergency rooms in Berlin. Dtsch Arztebl Int 2023; 120: 613–4 CrossRef MEDLINE PubMed Central VOLLTEXT
2.
Gans SL, Pols MA, Stoker J, Boermeester MA, on behalf of the expert steering group: Guideline for the diagnostic pathway in patients with ­acute abdominal pain. Dig Surg 2015; 32: 23–31 CrossRef MEDLINE
3.
Hoppe JM, Auer MK, Strüven A, Massberg S, Stremmel C: ChatGPT with GPT-4 outperforms emergency department physicians in diagnostic ­accuracy: retrospective analysis. J Med Internet Res 2024; 26: e56110 CrossRef MEDLINE PubMed Central
4.
Faqar-Uz-Zaman SF, Anantharajah L, Baumartz P, et al.: The diagnostic efficacy of an app-based diagnostic health care application in the ­emergency room: eRadaR-trial. A prospective, double-blinded, ­observational study. Ann Surg 2022; 276: 935–42 CrossRef MEDLINE
5.
Regierungskommission für eine moderne und bedarfsgerechte Krankenhausversorgung: Reform der Notfall- und Akutversorgung in Deutschland. 2023. www.bundesgesundheitsministerium.de/fileadmin/Dateien/3_Downloads/K/Krankenhausreform/Vierte_Stellungnahme_Regierungskommission_Notfall_ILS_und_INZ.pdf (last accessed on 19 February 2024).
Klinik und Poliklinik für Allgemein-, Viszeral-, Gefäß- und Transplantationschirurgie, Universitätsklinikum Bonn, Germany (Henn, Dohmen, Kalff, Matthaei) jonas.henn@ukbonn.de
Bonn Surgical Technology Center (BOSTER), Universitätsklinikum Bonn, Germany (Henn, Feodorovici, Dohmen, Kalff, Matthaei)
Klinik für Thoraxchirurgie, Universitätsklinikum Bonn, Germany (Feodorovici)
Abteilung für Klinische Akut- und Notfallmedizin, Universitätsklinikum Bonn, Germany (Gräff)
Direct comparison of sensitivity and specificity across evaluator groups.
Figure
Direct comparison of sensitivity and specificity across evaluator groups.
1.Helbig L, Möckel M, Fischer-Rosinsky A, Slagman A: Non-traumatic ­abdominal pain—a retrospective analysis of secondary data from 448 689 cases treated in two emergency rooms in Berlin. Dtsch Arztebl Int 2023; 120: 613–4 CrossRef MEDLINE PubMed Central VOLLTEXT
2.Gans SL, Pols MA, Stoker J, Boermeester MA, on behalf of the expert steering group: Guideline for the diagnostic pathway in patients with ­acute abdominal pain. Dig Surg 2015; 32: 23–31 CrossRef MEDLINE
3.Hoppe JM, Auer MK, Strüven A, Massberg S, Stremmel C: ChatGPT with GPT-4 outperforms emergency department physicians in diagnostic ­accuracy: retrospective analysis. J Med Internet Res 2024; 26: e56110 CrossRef MEDLINE PubMed Central
4.Faqar-Uz-Zaman SF, Anantharajah L, Baumartz P, et al.: The diagnostic efficacy of an app-based diagnostic health care application in the ­emergency room: eRadaR-trial. A prospective, double-blinded, ­observational study. Ann Surg 2022; 276: 935–42 CrossRef MEDLINE
5.Regierungskommission für eine moderne und bedarfsgerechte Krankenhausversorgung: Reform der Notfall- und Akutversorgung in Deutschland. 2023. www.bundesgesundheitsministerium.de/fileadmin/Dateien/3_Downloads/K/Krankenhausreform/Vierte_Stellungnahme_Regierungskommission_Notfall_ILS_und_INZ.pdf (last accessed on 19 February 2024).