Haoran Su, Liang You, Bailin Jiang, Zilong Wu, Guilan Kong, Yi Feng
{"title":"Machine Learning-based Early Detection of Intraoperative Anaphylaxis Among Patients with Hypotension Using Real-World Physiological Time Series Data.","authors":"Haoran Su, Liang You, Bailin Jiang, Zilong Wu, Guilan Kong, Yi Feng","doi":"10.1007/s10916-026-02452-8","DOIUrl":"https://doi.org/10.1007/s10916-026-02452-8","url":null,"abstract":"<p><strong>Background: </strong>Intraoperative anaphylaxis remains a rare yet fatal condition that has been a challenge due to its unpredictability. The unique characteristics of physiological parameter changes that precede or occur during the early stages of anaphylaxis may enable timely identification. We aimed to develop machine learning methods for the early detection of intraoperative anaphylaxis among patients with hypotension using real-world time-series physiological data.</p><p><strong>Methods: </strong>Physiological data, extracted from the EMRs at a 10-second sampling frequency, of patients undergoing surgeries at Peking University People's Hospital (January 1, 2011 - January 1, 2023) were analyzed. Three datasets (Datasets +1Min, +2Min, and +3Min) were constructed, each spanning 10 min before to 1, 2, and 3 min after hypotension, consisting of positive groups (intraoperative anaphylactic patients with hypotension) and negative groups (intraoperative non-anaphylactic patients with hypotension). Random forests, extreme gradient boosting, and categorical boosting (CatBoost) were employed to construct detection models. The model with the best performance was identified with a five-fold cross-validation.</p><p><strong>Results: </strong>Datasets +1Min, +2Min, and +3Min contained 49, 48, and 44 positive samples and 980, 960, and 880 negative samples, respectively. CatBoost models performed best on both Datasets +3Min and +2Min, achieving on Dataset +3Min an area under the receiver operator characteristic curve (AUROC) of 0.851 (± 0.082), an area under the precision-recall curve (AUPRC) of 0.431 (± 0.094), a sensitivity of 0.861 (± 0.210), and a specificity of 0.783 (± 0.164); corresponding values on Dataset +2Min were 0.823 (± 0.145), 0.365 (± 0.097), 0.711 (± 0.183), and 0.940 (± 0.065).</p><p><strong>Conclusions: </strong>This study suggests the feasibility of early detection of intraoperative anaphylaxis among patients with hypotension using physiological time-series data. The models could alert clinicians to suspected anaphylaxis just 2-3 min after hypotension, enabling early detection of intraoperative anaphylaxis.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-09-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148891454","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Clarifying Associations Between HIS Provider Characteristics and Hospital Digital Maturity: Provider-Level Inference, Self-Inclusion, and Market-Size Dependence.","authors":"Hao Lyu, Yaowen Hu, Shucai Fan","doi":"10.1007/s10916-026-02456-4","DOIUrl":"https://doi.org/10.1007/s10916-026-02456-4","url":null,"abstract":"","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-09-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148891507","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Hongjie Zhu, Xin Geng, Huayuan Zhou, Wei Guo, Yin Dai, Hao Zhang, Meng Dong, Hong Li, Xinlu Wang
{"title":"An AI-assisted Clinical Decision Support System for Green Classification of Cystocele on Dynamic Transperineal Ultrasound.","authors":"Hongjie Zhu, Xin Geng, Huayuan Zhou, Wei Guo, Yin Dai, Hao Zhang, Meng Dong, Hong Li, Xinlu Wang","doi":"10.1007/s10916-026-02453-7","DOIUrl":"https://doi.org/10.1007/s10916-026-02453-7","url":null,"abstract":"<p><p>Green classification of cystocele on dynamic transperineal ultrasound (TPUS) remains operator-dependent because it requires manual frame selection and landmark-based assessment of the Valsalva maneuver. We developed a workflow-oriented AI-assisted clinical decision support system for automated urethrovesical junction localization and dynamic Green classification and prospectively evaluated its standalone and reader-support performance. This diagnostic accuracy and reader study included 881 patients from a tertiary referral hospital, comprising a retrospective development cohort (n = 688) and an independent prospective test cohort (n = 193). A nested subset of 67 prospective patients was used for a reader study involving two junior and two intermediate radiologists under unaided and AI-assisted conditions. In the complete prospective test cohort, Green-AttGRU achieved a macro-averaged AUC of 0.939 (95% CI, 0.897-0.971) and an overall accuracy of 0.902 (95% CI, 0.860-0.943). In the reader study, overall accuracy increased from 0.761 to 0.821 without AI to 0.851-0.881 with AI, while macro-F1 increased from 0.660 to 0.777 to 0.820-0.860. Overall inter-reader agreement increased from a Fleiss' κ of 0.453 to 0.786, and pooled median interpretation time decreased from 26.7 s to 9.9 s. These findings support the preliminary feasibility of the system as a workflow-oriented decision-support tool for dynamic TPUS interpretation.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148874484","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Mattia Galanti, Filippo Battaglia, Giovanni Gugliandolo, Petter M Omland, Kristian B Nilsen, Sarah Breevoort, M Noreen Herring, Jerome Kurent, I-Hweii A Chen, Tuan Vu, Kyle Wuthrich, Nicola Donato, Mulugeta Gebregziabher, Giuseppe Campobello, Brian C Dean, Jonathan J Halford
{"title":"Evaluation of MPEG-4 AAC Audio Codec Using EEG and EMG Signals for Use with DICOM® Neurophysiology.","authors":"Mattia Galanti, Filippo Battaglia, Giovanni Gugliandolo, Petter M Omland, Kristian B Nilsen, Sarah Breevoort, M Noreen Herring, Jerome Kurent, I-Hweii A Chen, Tuan Vu, Kyle Wuthrich, Nicola Donato, Mulugeta Gebregziabher, Giuseppe Campobello, Brian C Dean, Jonathan J Halford","doi":"10.1007/s10916-026-02451-9","DOIUrl":"10.1007/s10916-026-02451-9","url":null,"abstract":"<p><p>Modern neurophysiological recording technologies are producing large volumes of data. This study investigates the use of the Moving Picture Experts Group version 4 Advanced Audio Coding (MPEG-4 AAC) for compressing electroencephalography (EEG) and electromyography (EMG) signals. All EEG signals contained either seizures or interictal epileptiform discharges (IEDs), which are nonstationary patterns characterized by sharp transients. Differences between pairs of original single-channel signals and compressed/decompressed signals were explored using the Percentage Root Mean Square Difference (PRD), power spectral density (PSD) features, and clinical expert opinion of waveform morphology. PRD was used as the primary error measure and PSD features were explored across different frequency bands. The results showed substantial decline in signal quality based on human expert opinion for compressed then reconstructed EEG signals with PRD error greater than 15% and EMG signals with PRD greater than 1%. PSD analysis showed significant changes in the gamma and beta bands for both EEG and EMG signals. This study provides further evidence that standard audio codecs which use psychoacoustic models, such as MPEG-4 AAC, do not sufficiently support compression of neurophysiology waveforms.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-08-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518441/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148828950","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Clinical AI Generators and Reviewers must be Tested Together.","authors":"Vera Sorin, Eyal Klang","doi":"10.1007/s10916-026-02454-6","DOIUrl":"https://doi.org/10.1007/s10916-026-02454-6","url":null,"abstract":"","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-08-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148818766","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Andreas Praschk, Valentin Fischill-Neudeck, Thomas Caspari, Hans-Peter Wiesinger
{"title":"Beyond Accuracy: A Mixed-Methods Audit of Chain-of-Thought Failures in LLM-Based COVID-19 Vaccine Stance Detection.","authors":"Andreas Praschk, Valentin Fischill-Neudeck, Thomas Caspari, Hans-Peter Wiesinger","doi":"10.1007/s10916-026-02449-3","DOIUrl":"10.1007/s10916-026-02449-3","url":null,"abstract":"<p><p>This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring's content analysis guided by the FUTURE-AI framework. At each model's best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95% CI 0.008-0.032). Under the reasoning-intensive settings, CoT availability differed: Gemini returned a reasoning summary for all tweets, whereas o4-mini did so for 64.7%. Among 1,981 tweets with CoTs from both models, 295 (14.9%) were dual-errors; in 88.8%, both models produced the same wrong label, suggesting shared failure modes. Qualitatively, both models showed the same errors: target confusion (policy vs. vaccine), literal readings of sarcasm, and label-rationale mismatches, recurring across models despite their markedly different CoT lengths. Reasoning LLMs can therefore classify stance accurately, but their readiness for transparent public health applications depends on whether a CoT is available at all and whether it is coherent with the label it accompanies (label-rationale coherence). CoT availability, label-rationale coherence, and safeguards against systematic reasoning failures offer candidate explainability-readiness metrics, alongside accuracy, for trustworthy digital epidemiology.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-08-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13452790/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148697557","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Jonas Backes, Alexander Geissler, Jonas Subelack, David Ehlig
{"title":"Does System Choice Matter? Influence of Hospital Information Systems on Digital Maturity.","authors":"Jonas Backes, Alexander Geissler, Jonas Subelack, David Ehlig","doi":"10.1007/s10916-026-02450-w","DOIUrl":"10.1007/s10916-026-02450-w","url":null,"abstract":"<p><p>Hospitals worldwide invest heavily in digital infrastructure to improve efficiency, safety, and patient-centeredness. Hospital Information Systems (HIS) form the backbone of these efforts. However, the HIS provider market remains fragmented and heterogeneous, thus potentially influencing hospitals' ability to achieve digital transformation. Empirical evidence on how HIS choice relates to digital maturity remains limited. Using data from more than 1,600 hospitals participating in the German-wide DigitalRadar (DR) project (i.e., covering about 90% of all hospitals), we examined the relationship between HIS provider characteristics and hospitals' digital maturity. The DR dataset comprises 234 standardized items per hospital and provides a maturity score from 0 to 100 for each hospital. Four HIS provider features (i.e., module utilization, integration ratio, external provider variation, and maximum hospital size coverage) were derived and used for k-means clustering. Multivariate linear regressions, controlling for hospital characteristics, quantified associations with digital maturity. K-means clustering identified four clusters of HIS providers. Higher module utilization (β = 6.87, p < 0.05) and integration ratios (β = 14.35, p < 0.01) were consistently associated with greater digital maturity. External provider variation and maximum hospital size coverage were not significantly associated with overall digital maturity. The choice of HIS provider is associated with hospitals' digital maturity, particularly with respect to broader module coverage and stronger integration. Providers with low module utilization and limited integration are associated with lower maturity. Our findings highlight the need for strategic HIS procurement, targeted system/organizational improvements, and policy incentives to accelerate digital transformation.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-08-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13452863/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148697642","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Temporal Variability in Diagnosis Code Distributions Across Extraction Time Points in a Multicenter Integrated EHR Database: A Snapshot Comparison Study.","authors":"Kyunghee Lee, Chihiro Kumagai, Masato Komuro, Kengo Miyo, Hiroyuki Hoshimoto","doi":"10.1007/s10916-026-02447-5","DOIUrl":"10.1007/s10916-026-02447-5","url":null,"abstract":"<p><p>EHR data are widely used in clinical research; however, their temporal stability remains poorly understood. This study aimed to evaluate temporal variability in EHR data arising from differences in extraction time points. We conducted a retrospective, descriptive, observational study using diagnostic data from a multicenter integrated EHR database in Japan. The target period was 2022, and diagnosis records with a start date within this period were extracted at three time points: October 2023, October 2024, and October 2025. Diagnoses were mapped to the three-character level of the International Classification of Diseases, 10th Revision (ICD-10). Changes in diagnosis code distributions were quantified using the Jensen-Shannon distance (JSD). Changes in record counts per ICD-10 code between 2023 and 2025 were also evaluated. Five facilities were included in the analysis. In the integrated dataset, variations in total records and unique patients were minimal (total records: +0.06% in 2024 and + 0.05% in 2025; unique patients: +0.06% in 2024 and + 0.02% in 2025). JSD values for ICD-10 distributions were non-zero across all facilities, indicating that diagnosis code distributions differed across extraction time points even for the same target period. Across all facilities excluding Facility C, JSD was 0.0022 in 2024 and 0.0026 in 2025 relative to the 2023 baseline. In contrast, Facility C exhibited a markedly larger JSD in 2024 (0.0310), which persisted at a similar level in 2025 (0.0317). At the code level, H61 showed the largest absolute change (- 266 records; -19.98%). Overall, acute conditions tended to decrease, whereas chronic conditions increased across extraction time points. EHR data exhibit temporal variability across extraction time points, even for the same target period. These findings highlight the importance of documenting extraction timing and dataset version, and suggest the need for study designs that explicitly account for such variability.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-08-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13442496/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148678932","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"When Useful Clinical AI Exceeds Meaningful Oversight.","authors":"Jonas Ver Berne, Reinhilde Jacobs","doi":"10.1007/s10916-026-02446-6","DOIUrl":"10.1007/s10916-026-02446-6","url":null,"abstract":"<p><p>Human oversight is widely invoked as a safeguard in clinical AI, yet it is often treated as a single, stable property. In practice, clinicians remain responsible for AI-assisted decisions without remaining able to meaningfully review the basis of the output. This tension becomes sharper as systems become multimodal, agentic, and cognitively ambitious. We propose a distinction between cognitive and normative oversight and discuss its implications for AI evaluation and deployment.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148648837","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Automated Multi-Label Adverse drug Reaction Extraction from Malaysia's National Pharmacovigilance Narratives.","authors":"Sze Gee Lim, Maizatul Akmar Ismail","doi":"10.1007/s10916-026-02448-4","DOIUrl":"https://doi.org/10.1007/s10916-026-02448-4","url":null,"abstract":"<p><p>Extracting multiple adverse drug reaction (ADR) terms from unstructured narratives remains challenging, particularly under severe label imbalance that limits the detection of rare ADRs. This study aimed to develop and evaluate a multi-label natural language processing framework for automated ADR extraction within Malaysia's national pharmacovigilance reporting system. We evaluated classical machine learning and transformer-based models within a common framework, incorporating domain-specific preprocessing and imbalance-aware optimization strategies to improve the detection of rare ADRs. The framework was applied to 28,980 real-world ADR narratives annotated with 80 Medical Dictionary for Regulatory Activities (MedDRA) Preferred Terms (PTs) from the Skin and Subcutaneous Tissue Disorders System Organ Class (SOC). Performance was evaluated using micro-F1, macro-F1, precision, recall, and label coverage using a predefined 80/20 train-test evaluation strategy, complemented by pharmacist pilot review. Transformer-based models achieved the strongest performance, with the augmented model attaining a micro-F1 of 0.88, macro-F1 of 0.55, recall of 0.88, and the highest label coverage, correctly predicting 60 of 80 PTs (75%). However, a 70/10/20 sensitivity analysis on the improved transformer model yielded lower micro-F1 and precision, highlighting the impact of further data partitioning on model performance. Improved classical models also showed substantial gains over their baseline models, particularly in micro-F1, macro-F1, recall, and label coverage, consistent with improved detection of rare ADRs. During the pilot review, pharmacists accepted the model-preselected PTs without modification in 19 of 25 narratives (76%), while additional PTs were added in the remaining cases. Although the study was limited to a single SOC and one national pharmacovigilance database, the findings demonstrate that optimization strategies, domain-specific feature enrichment, and targeted augmentation can substantially improve multi-label ADR classification under severe imbalance. These findings support further evaluation of AI-assisted pharmacovigilance with pharmacist oversight in real-world practice.</p>","PeriodicalId":16338,"journal":{"name":"Journal of Medical Systems","volume":"50 1","pages":""},"PeriodicalIF":8.8,"publicationDate":"2026-07-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148630690","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}