A Machine Learning Approach for Identifying People With Neuroinfectious Diseases in Electronic Health Records: Algorithm Development and Validation.

Arjun Singh , Shadi Sartipi , Haoqi Sun , Rebecca Milde , Niels Turley , Carson Quinn , G Kyle Harrold , Rebecca L Gillani , Sarah E Turbett , Sudeshna Das , Sahar Zafar , Marta Fernandes , M Brandon Westover , Shibani S Mukerji

JMIR Med Inform

Department of Neurology, Massachusetts General Hospital, 55 Fruit St, Wang ACC 835, Boston, MA, 02114, United States, 1 2163379887.

Published: August 2025

Category Ranking

98%

Total Visits

921

Avg Visit Duration

2 minutes

Citations

Background: Identifying neuroinfectious disease (NID) cases using International Classification of Diseases billing codes is often imprecise, while manual chart reviews are labor-intensive. Machine learning models can leverage unstructured electronic health records to detect subtle NID indicators, process large data volumes efficiently, and reduce misclassification. While accurate NID classification is needed for research and clinical decision support, using unstructured notes for this purpose remains underexplored.

Objective: The objective of this study is to develop and validate a machine learning model to identify NIDs from unstructured patient notes.

Methods: Clinical notes from patients who had undergone lumbar puncture were obtained using the electronic health record of an academic hospital network (Mass General Brigham [MGB]), with half associated with NID-related diagnostic codes. Ground truth was established by chart review with 6 NID-expert physicians. NID keywords were generated with regular expressions, and extracted texts were converted into bag-of-words representations using n-grams (n=1, 2, 3). Notes were randomly split into training (80%), 2400 notes out of 3000, and hold-out testing (20%), 600 notes out of 3000, sets. Feature selection was performed using logistic regression with L1 regularization. An extreme gradient boosting (XGBoost) model classified NID cases, and performance was evaluated using the area under the receiver operating curve (AUROC) and the precision-recall curve (AUPRC). The performance of the natural language processing (NLP) model was contrasted with the Llama 3.2 auto-regressive model on the MGB test set. The NLP model was additionally validated on external data from an independent hospital (Beth Israel Deaconess Medical Center [BIDMC]).

Results: This study included 3000 patient notes from MGB from January 22, 2010, to September 21, 2023. Of 1284 initial n-gram features, 342 were selected, with the most significant features being "meningitis," "ventriculitis," and "meningoencephalitis." The XGBoost model achieved an AUROC of 0.98 (95% CI 0.96-0.99) and AUPRC of 0.89 (95% CI 0.83-0.94) on MGB test data. In comparison, NID identification using International Classification of Diseases billing codes showed high sensitivity (0.97) but poor specificity (0.59), overestimating NID cases. Llama 3.2 improved specificity (0.94) but had low sensitivity (0.64) and an AUROC of 0.80. In contrast, our NLP model balanced specificity (0.96) and sensitivity (0.84), outperforming both methods in accuracy and reliability on MGB data. When tested on external data from BIDMC, the NLP model maintained an AUROC of 0.98 (95% CI 0.96-0.99), with an AUPRC of 0.78 (95% CI 0.66-0.89).

Conclusions: The NLP model accurately identifies NID cases from clinical notes. Validated across 2 independent hospital datasets, the model demonstrates feasibility for large-scale NID research and cohort generation. With further external validation, our results could be more generalizable to other institutions.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC12396800	PMC
http://dx.doi.org/10.2196/63157	DOI Listing

Publication Analysis

Top Keywords

nlp model

nid cases

machine learning

electronic health

model

nid

health records

international classification

classification diseases

diseases billing

Similar Publications

Psychology's Questionable Research Fundamentals (QRFs): Key problems in quantitative psychology and psychological measurement beyond Questionable Research Practices (QRPs).

Front Psychol

August 2025

Faculty of Information Technology, Monash University, Melbourne, VIC, Australia.

Jana Uher , Jan Ketil Arnulf , Paul T Barrett , Moritz Heene , Jörg-Henrik Heine

Psychology's crises (e.g., replicability, generalisability) are currently believed to derive from Questionable Research Practices (QRPs), thus scientific misconduct.

View Article and Find Full Text PDF

Similar Publications

A Pure Transformer Pretraining Framework on Text-attributed Graphs.

Proc Mach Learn Res

November 2024

Michigan State University.

Yu Song , Haitao Mao , Jiachen Xiao , Jingzhe Liu , Zhikai Chen

Pretraining plays a pivotal role in acquiring generalized knowledge from large-scale data, achieving remarkable successes as evidenced by large models in CV and NLP. However, progress in the graph domain remains limited due to fundamental challenges represented by feature heterogeneity and structural heterogeneity. Recent efforts have been made to address feature heterogeneity via Large Language Models (LLMs) on text-attributed graphs (TAGs) by generating fixed-length text representations as node features.

View Article and Find Full Text PDF

Similar Publications

[The Dawn of Generative AI in Medicine: Empathy Through Emulation].

Dtsch Med Wochenschr

September 2025

Klinik für Kardiologie, Angiologie und Pneumologie, Institut für Cardiomyopathien Heidelberg, Universitätsklinikum Heidelberg, Heidelberg, Deutschland.

Ali Amr , Benjamin Meder

Rapid advancements in Artificial Intelligence (AI) have significantly impacted multiple sectors of our society, including healthcare. While conventional AI has been instrumental in solving mainly image recognition tasks and thereby adding in well-defined situations such as supporting diagnostic imaging, the emergence of generative AI is impacting on one of the main professional competences: doctor-patient interaction.A convergence of natural language processing (NLP) and generative AI is exemplified by intelligent chatbots like ChatGPT.

View Article and Find Full Text PDF

Similar Publications

Integrating rule-based NLP and large language models for statin information extraction from clinical notes.

Int J Med Inform

September 2025

Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, USA; Department of Medicine, Vanderbilt University Medical Center, Nashville, TN, USA. Electronic address:

Siru Liu , Allison B McCoy , Qingyu Chen , Adam Wright

Background: Identifying patient-specific barriers to statin therapy, such as intolerance or deferral, from clinical notes is a major challenge for improving cardiovascular care. Automating this process could enable targeted interventions and improve clinical decision support (CDS).

Objective: To develop and evaluate a novel hybrid artificial intelligence (AI) framework for accurately and efficiently extracting information on statin therapy barriers from large volumes of clinical notes.

View Article and Find Full Text PDF

Similar Publications

Analyzing Depression in College Students Using NLP and Transformer Models: Implications for Career and Educational Counseling.

Brain Behav

September 2025

Department of Electrical and Computer Engineering, University of Tabriz, Tabriz, Iran.

Qiuxia Wan , Yue Pan , Sonia Zakeri

Purpose: Depression among college students is a growing concern that negatively affects academic performance, emotional well-being, and career planning. Existing diagnostic methods are often slow, subjective, and inaccessible, underscoring the need for automated systems that can detect depressive symptoms through digital behavior, particularly on social media platforms.

Method: This study proposes a novel natural language processing (NLP) framework that combines a RoBERTa-based Transformer with gated recurrent unit (GRU) layers and multimodal embeddings.

View Article and Find Full Text PDF

Similar Publications