主页 文献库文献详情
PMID: 41890367 已发表 · epublish 英语

Scaling sensor metadata extraction for exposure health using LLMs.

Exposome ·第 6 卷 ·第 1 期

Shah-Mohammadi F, Im S, Facelli JC, Cummins MR, Gouripeddi R

摘要

The rapid evolution and diversity of sensor technologies, coupled with inconsistencies in how sensor metadata is reported across formats and sources, present significant challenges for generating exposomes and exposure health research. Despite the development of standardized metadata schemas, the process of extracting sensor metadata from unstructured sources remains largely manual and unscalable. To address this bottleneck, we developed and evaluated a large language model (LLM)-based pipeline for automating sensor metadata extraction and harmonization from publicly available exposure health literature. Using GPT-4 in a zero-shot setting, we constructed a pipeline that parses full-text PDFs to extract metadata and harmonizes output into structured formats. Our automated pipeline achieved substantial efficiency gains in completing extractions much faster than manual review and demonstrated strong performance with 88.0% accuracy, 88.0% precision, 93.0% recall, and an F1-score of 90.0%. This study demonstrates the feasibility and scalability of leveraging LLMs to automate sensor metadata extraction for exposure health, reducing manual burden while enhancing metadata completeness and consistency. Our findings support the integration of LLM-driven pipelines into exposure health informatics platforms.

关键词
GPT exposure health information extraction metadata sensor
文献信息
期刊
Exposome
期刊简称
Exposome
ISSN
2635-2265
语言
英语
国家/地区
England
NLM ID
9918317685206676
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: product@genelibs.com