Meyer, Annika ORCID: 0000-0002-8411-8799, Schömig, Edgar and Streichert, Thomas ORCID: 0000-0002-6588-720X (2025). ChatGPT and reference intervals: a comparative analysis of repeatability in GPT-3.5 Turbo, GPT-4, and GPT-4o. Frontiers in Artificial Intelligence, 8. pp. 1-10. Frontiers Media. ISSN 2624-8212

[thumbnail of frai-8-1681979.pdf] PDF
frai-8-1681979.pdf
Bereitstellung unter der CC-Lizenz: Creative Commons Attribution.

Download (3MB)
Identification Number:10.3389/frai.2025.1681979

Abstract

[Artikel-Nr.: 1681979] Background: Large language models such as ChatGPT hold promise as rapid “curbside consultation” tools in laboratory medicine. However, their ability to generate consistent and clinically reliable reference intervals—particularly in the absence of contextual clinical information—remains uncertain. Method: This cross-sectional study evaluated whether three versions of ChatGPT (GPT-3.5-Turbo, GPT-4, GPT-4o) maintain repeatable reference-interval outputs when the prompt intentionally omits the interval, using reference interval variability as a stress-test for model consistency. Standardized prompts were submitted through 726,000 chatbot requests. A total of 246,842 reference intervals across 47 laboratory parameters were then analyzed for consistency using the coefficient of variation (CV) and regression models. Results: On average, the chatbots exhibited a CV of 26.50% (IQR: 7.35–129.01%) for the lower limit and 15.82% (IQR: 4.50–45.30%) for the upper limit upon repetition. GPT-4 and GPT-4o demonstrated significantly lower CVs compared to GPT-3.5-Turbo. Reference intervals for poorly standardized parameters were particularly inconsistent across lower ( β : 0.6; 95% CI: 0.35 to 0.86; p < 0.001) and upper limit (β: 0.5; 95% CI: 0.28 to 0.71; p < 0.001), while unit expressions also showed variability. Conclusion: While the newer ChatGPT versions tested demonstrate improved repeatability, diagnostically unacceptable variability persists, particularly for poorly standardized analytes. Mitigating this requires thoughtful prompt design (e.g., mandatory inclusion of reference intervals), global harmonization of laboratory standards, further model refinement, and robust regulatory oversight. Until then, AI chatbots should be restricted to professional use and trained to refuse laboratory interpretation when reference intervals are not provided by the user.

Item Type: Article
Creators:
Creators
Email
ORCID
ORCID Put Code
Meyer, Annika
UNSPECIFIED
UNSPECIFIED
Schömig, Edgar
UNSPECIFIED
UNSPECIFIED
UNSPECIFIED
Streichert, Thomas
UNSPECIFIED
UNSPECIFIED
URN: urn:nbn:de:hbz:38-812219
Identification Number: 10.3389/frai.2025.1681979
Journal or Publication Title: Frontiers in Artificial Intelligence
Volume: 8
Page Range: pp. 1-10
Number of Pages: 10
Date: 12 December 2025
Publisher: Frontiers Media
ISSN: 2624-8212
Language: English
Faculty: Faculty of Medicine
Divisions: Faculty of Medicine > Anästhesiologie und Operative Intensivmedizin > Klinik für Anästhesiologie und Operative Intensivmedizin
Faculty of Medicine > Klinische Chemie > Institut für Klinische Chemie
Faculty of Medicine > Pharmakologie > Instítut für Pharmakologie
Subjects: Medical sciences Medicine
Uncontrolled Keywords:
Keywords
Language
chatbot ; ChatGPT ; reference interval ; repeatability ; consistency ; large language model
English
['eprint_fieldname_oa_funders' not defined]: Publikationsfonds UzK
Refereed: Yes
URI: http://kups.ub.uni-koeln.de/id/eprint/81221

Downloads

Downloads per month over past year

Altmetric

Export

Actions (login required)

View Item View Item