ECG-based detection of occlusion myocardial infarction: a dedicated deep neural network versus multimodal large language models and physicians — a retrospective diagnostic accuracy study

  • Powerful Medical
  • August 7, 2026
  • 3 min to read
ECG detection of occlusion MI — dedicated deep neural network versus multimodal LLMs and physicians

Overview

Emergency ECG interpretation for acute coronary occlusion is shifting from the STEMI paradigm toward the broader occlusion myocardial infarction (OMI) concept. This retrospective diagnostic accuracy study compared the OMI-trained PMcardio Queen of Hearts (QoH) AI-ECG model against two general-purpose multimodal LLMs (ChatGPT 5.2, Gemini 3 Pro) and ten emergency physicians (five specialists, five residents) on 36 ECGs from patients referred for emergent coronary angiography. QoH achieved the highest discrimination (AUC 0.96) and sensitivity (95.8%), significantly outperforming physicians and both LLMs, while the LLMs showed only fair-to-moderate answer consistency across repeated queries and, for Gemini 3 Pro, marked overconfidence on incorrect diagnoses.

Key findings

  • QoH reached the highest discrimination of any interpreter (AUC 0.96, 95% CI 0.91–1.00) and 95.8% sensitivity, missing only 1 of 24 confirmed coronary occlusions.
  • QoH's sensitivity significantly exceeded EM specialists (66.7%, p=0.016), ChatGPT 5.2 (54.2%, p=0.006), Gemini 3 Pro (58.3%, p=0.004), and EM residents (p=0.008).
  • Across 5 repeated queries per ECG, the LLMs' diagnoses were only fair-to-moderately consistent (Fleiss' κ 0.24–0.49), and Gemini 3 Pro stayed at 95% median confidence even when wrong (Brier score 0.40, worse than a random-guess baseline of 0.25).
  • Gemini 3 Pro had the lowest accuracy (55.6%) and AUC (0.62) of any interpreter tested, including EM residents (58.3% accuracy).

Published in: BMC Emergency Medicine
Published on: 07 August 2026

Background

Accurate ECG interpretation for acute coronary occlusion is a critical, time-sensitive task in emergency care, and the field is shifting from the STEMI paradigm to the broader occlusion myocardial infarction (OMI) concept. This retrospective diagnostic accuracy study compared a dedicated occlusion-detection deep neural network (PMcardio Queen of Hearts, QoH) against two general-purpose multimodal large language models (ChatGPT 5.2, Gemini 3 Pro) and emergency physicians for ECG-based OMI detection.

Methods

The authors evaluated 36 twelve-lead ECGs (24 angiographically confirmed OMI, 12 non-OMI) drawn from patients referred for emergent coronary angiography (McCabe et al. dataset). Each ECG was interpreted by QoH (queried once, deterministic), by ChatGPT 5.2 and Gemini 3 Pro (each queried 5 times per ECG in separate sessions to test response consistency), and by 5 emergency medicine specialists and 5 emergency medicine residents.

Results

QoH showed the strongest discrimination (AUC 0.96, 95% CI 0.91–1.00), with 86.1% accuracy, 95.8% sensitivity, and 66.7% specificity. ChatGPT 5.2 reached an AUC of 0.81 (66.7% accuracy, 54.2% sensitivity, 91.7% specificity), while Gemini 3 Pro had the weakest performance (AUC 0.62, 55.6% accuracy, 58.3% sensitivity, 50.0% specificity). EM specialists achieved 75.0% accuracy, 66.7% sensitivity, and 91.7% specificity — the best result among physicians — while EM residents reached 58.3% accuracy. QoH's sensitivity was significantly higher than every other interpreter, including EM specialists (p=0.016), ChatGPT 5.2 (p=0.006), Gemini 3 Pro (p=0.004), and residents (p=0.008); the AUC gap versus ChatGPT (Δ0.15) and the accuracy gap versus EM specialists were numerically large but did not reach significance in this underpowered 36-ECG sample. Across the 5 repeated runs, LLM responses were only fair-to-moderately consistent (Fleiss' κ 0.24–0.49). Gemini 3 Pro was poorly calibrated, reporting a median confidence of 95% even on incorrect diagnoses (Brier score 0.40, worse than the random-guess baseline of 0.25), whereas ChatGPT 5.2 was better calibrated (Brier 0.230) with a small but real confidence drop on wrong answers.

Conclusion

The dedicated occlusion-detection model showed the strongest and most clinically favorable profile, particularly for sensitivity — the most dangerous error in this setting is a missed coronary occlusion. General-purpose multimodal LLMs are not currently safe for autonomous OMI diagnosis and, at most, could serve as an adjunct or second-reader tool pending prospective validation.

Authors: Emin Hüseyin Akar, Kamil Kokulu, Ekrem Taha Sert

Share this article

  • Share on Facebook
  • Share on X
  • Share on Linkedin
  • Share via e-mail
  • Copy the link
About PMcardio logo

About PMcardio

PMcardio is the market leader in AI-powered diagnostics, addressing the world’s leading cause of death – cardiovascular diseases. The innovative clinical assistant empowers healthcare professionals to detect up to 40 cardiovascular diseases. In the form of a smartphone application, the certified Class IIb medical device interprets any 12-lead ECG image in under 5 seconds to provide accurate diagnoses and individualized treatment recommendations tailored to each patient.

About Product

About Powerful Medical  logo

About Powerful Medical

Powerful Medical leads one of the most important shifts in modern medicine by augmenting human-made clinical decisions with artificial intelligence. Our primary focus is on cardiovascular diseases, the world’s leading cause of death.

Established in 2017, Powerful Medical has embarked on a mission to revolutionize the diagnosis and treatment of cardiovascular diseases. We are a medical company backed by 28 world-class cardiologists and led by our expert Scientific Board with decades of experience in daily patient care, clinical research, and medical devices. The results of our research are implemented, developed, certified, and brought to market by our 50+ strong interdisciplinary team of physicians, data scientists, AI experts, software engineers, regulatory specialists, and commercial teams.

About Us