Meyer, Annika
ORCID: 0000-0002-8411-8799, Schömig, Edgar and Streichert, Thomas
ORCID: 0000-0002-6588-720X
(2025).
ChatGPT and reference intervals: a comparative analysis of repeatability in GPT-3.5 Turbo, GPT-4, and GPT-4o.
Frontiers in Artificial Intelligence, 8.
pp. 1-10.
Frontiers Media.
ISSN 2624-8212
|
PDF
frai-8-1681979.pdf Bereitstellung unter der CC-Lizenz: Creative Commons Attribution. Download (3MB) |
Abstract
[Artikel-Nr.: 1681979] Background: Large language models such as ChatGPT hold promise as rapid “curbside consultation” tools in laboratory medicine. However, their ability to generate consistent and clinically reliable reference intervals—particularly in the absence of contextual clinical information—remains uncertain. Method: This cross-sectional study evaluated whether three versions of ChatGPT (GPT-3.5-Turbo, GPT-4, GPT-4o) maintain repeatable reference-interval outputs when the prompt intentionally omits the interval, using reference interval variability as a stress-test for model consistency. Standardized prompts were submitted through 726,000 chatbot requests. A total of 246,842 reference intervals across 47 laboratory parameters were then analyzed for consistency using the coefficient of variation (CV) and regression models. Results: On average, the chatbots exhibited a CV of 26.50% (IQR: 7.35–129.01%) for the lower limit and 15.82% (IQR: 4.50–45.30%) for the upper limit upon repetition. GPT-4 and GPT-4o demonstrated significantly lower CVs compared to GPT-3.5-Turbo. Reference intervals for poorly standardized parameters were particularly inconsistent across lower ( β : 0.6; 95% CI: 0.35 to 0.86; p < 0.001) and upper limit (β: 0.5; 95% CI: 0.28 to 0.71; p < 0.001), while unit expressions also showed variability. Conclusion: While the newer ChatGPT versions tested demonstrate improved repeatability, diagnostically unacceptable variability persists, particularly for poorly standardized analytes. Mitigating this requires thoughtful prompt design (e.g., mandatory inclusion of reference intervals), global harmonization of laboratory standards, further model refinement, and robust regulatory oversight. Until then, AI chatbots should be restricted to professional use and trained to refuse laboratory interpretation when reference intervals are not provided by the user.
| Item Type: | Article |
| Creators: | Creators Email ORCID ORCID Put Code Schömig, Edgar UNSPECIFIED UNSPECIFIED UNSPECIFIED |
| URN: | urn:nbn:de:hbz:38-812219 |
| Identification Number: | 10.3389/frai.2025.1681979 |
| Journal or Publication Title: | Frontiers in Artificial Intelligence |
| Volume: | 8 |
| Page Range: | pp. 1-10 |
| Number of Pages: | 10 |
| Date: | 12 December 2025 |
| Publisher: | Frontiers Media |
| ISSN: | 2624-8212 |
| Language: | English |
| Faculty: | Faculty of Medicine |
| Divisions: | Faculty of Medicine > Anästhesiologie und Operative Intensivmedizin > Klinik für Anästhesiologie und Operative Intensivmedizin Faculty of Medicine > Klinische Chemie > Institut für Klinische Chemie Faculty of Medicine > Pharmakologie > Instítut für Pharmakologie |
| Subjects: | Medical sciences Medicine |
| Uncontrolled Keywords: | Keywords Language chatbot ; ChatGPT ; reference interval ; repeatability ; consistency ; large language model English |
| ['eprint_fieldname_oa_funders' not defined]: | Publikationsfonds UzK |
| Refereed: | Yes |
| URI: | http://kups.ub.uni-koeln.de/id/eprint/81221 |
Downloads
Downloads per month over past year
Altmetric
Export
Actions (login required)
![]() |
View Item |
https://orcid.org/0000-0002-8411-8799