主页 文献库文献详情
PMID: 42428234 已发表 · epublish 英语

Evaluation of Large Language Models in the Clinical Management of Patients With Upper Gastrointestinal Bleeding: Insights From Real-World Patient Data.

DEN open ·第 7 卷 ·第 1 期 ·2027-04-00

Rajabnia M, Davoodi F, Hajisafarali M, Asadinia S, Shiri M, Shirzad F, Saeedi P, Mohammadi M

摘要

To evaluate large language models (LLMs) for pre-endoscopy (PE) risk stratification and prediction of endoscopic findings in upper gastrointestinal bleeding (UGIB), and compare their performance with clinical risk scores, conventional machine learning (ML) models, and hybrid approaches. This multicenter retrospective study included 384 patients with UGIB who underwent endoscopy. Five LLMs (GPT-5, Gemini-2.5-Flash, Llama 4, Grok, and DeepSeek R1) were tested using structured zero-shot prompts based on PE clinical and laboratory data. Their performance in identifying high-risk patients was compared with the Glasgow-Blatchford Score (GBS), AIMS65, PE Rockall score, and ML models. Hybrid LLM-score models were also assessed. Two gastroenterologists evaluated the quality of LLM-generated justifications. GBS showed the best discriminative performance among clinical scores (area under the receiver operating characteristic curve [AUROC] 0.73). Among LLMs, GPT-5 achieved the highest accuracy (0.66), while Grok showed the best-balanced performance (0.59; F1 0.47). Gemini-2.5-Flash had the highest sensitivity (0.89) but low specificity (0.21). Endoscopic prediction performance was modest, with Gemini-2.5-Flash achieving the highest exact-match accuracy (0.34) and micro-F1 (0.38). Hybrid models improved performance over standalone LLMs but did not outperform GBS alone (best: GBS+GPT-5, AUROC 0.670. LLMs showed higher numerical performance than conventional ML models, but a statistical comparison was not possible due to unavailable instance-level data. Grok received the highest human evaluation score for explanation quality. LLMs showed moderate performance in UGIB risk stratification and endoscopic prediction but were inferior to clinical scores, especially GBS. Hybrid models modestly improved over standalone LLMs but not GBS, supporting their use as adjunct tools rather than clinical decision-support systems. N/A.

关键词
artificial intelligence clinical decision support endoscopy gastrointestinal bleeding large language models
文献信息
期刊
DEN open
期刊简称
DEN Open
ISSN
2692-4609
发表日期
2027-04-00
语言
英语
国家/地区
Australia
NLM ID
9918317682706676
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: product@genelibs.com