◆ AI & Medical AI Daily
–
Monday · Clinical Applications & Research

Clinical AI's week produced no new trial — it produced three sets of accounts. What actually took in-hospital mortality from 23.1% to 18.6% was a deterioration score with no LLM in it; the GPT-4o randomized trial stopped at aOR 0.77, P = 0.13

This issue has to open with an unflattering sentence: in the past 72 hours, no new large prospective trial of clinical AI landed. The two freshest documents here are both retrospective accounting — the NEJM AI study brought back into discussion on October 2, which cut in-hospital mortality among high-risk inpatients from 23.1% to 18.6% across 11 acute care hospitals and 23,132 encounters (adjusted aOR 0.82, 95% CI 0.74–0.91), and the 2026 Q3 quarterly review published the same day, which tallies a full quarter of trials and leaves one line standing: the workflow is now the model. Set those beside Nature Medicine's GPT-4o cluster-randomized trial in Kenya (9,691 patients; primary endpoint 2.2% versus 2.0%, aOR 0.77, 95% CI 0.55–1.08, P = 0.13) and today's conclusion is hard to walk around: what moves patient outcomes is the workflow, not the size of the model.

01 — Top Stories

Today's eight are deliberately ordered with patient outcomes first and scores last — because that is exactly where this week's difference lies
Inpatient mortality RWJBarnabas HealthEpic Deterioration IndexNEJM AI10/02

Wire Epic's deterioration index straight to the rapid response team, and high-risk inpatient mortality across 11 hospitals falls from 23.1% to 18.6%

What

RWJBarnabas Health and Rutgers Robert Wood Johnson Medical School routed real-time alerts from the Epic Deterioration Index (EDI) directly to the rapid response team across 11 acute care hospitals, instead of parking them on a nursing-station dashboard. The study appeared in NEJM AI (DOI 10.1056/AIoa2500973), led by Thomas Nahass, the system's VP of health informatics: 10,803 encounters pre-implementation and 12,329 post, 23,132 in total. On October 2 a 2 Minute Medicine appraisal pulled it back into view with the full statistics: unadjusted mortality 23.1% → 18.6% (absolute difference −4.5 percentage points, 95% CI −5.6 to −3.5); adjusted aOR 0.82 (95% CI 0.74–0.91); RRT activation OR 1.74 (95% CI 1.61–1.88), with RRT evaluations rising from 25.3% to 37.5% while ICU transfers did not rise proportionally.

Why it matters

This is one of the year's few clinical-AI studies that moved mortality rather than accuracy — and nothing in it is new. The EDI is a gradient-boosted score built into Epic, driven mostly by structured physiological values, with no LLM anywhere in it. What changed was the workflow: the alert no longer waits for someone to notice it, it summons a team that goes to the bedside. Nahass himself keeps the causal claim modest: "Our goal was to identify patients earlier, before they reached a point where intervention becomes much more difficult. The deterioration index gives us an earlier point in time."

Discount this

This is a prospective before-and-after cohort, not a randomized trial: the pre- and post-periods have unequal seasonal distributions, two of the hospitals implemented incompletely, and the authors flag generalizability as uncertain. Two figures deserve more weight still. Only 46.6% of post-implementation encounters actually generated a notification; and in the EDI 55–59 band that sits just across the threshold, the adjusted aOR is 0.86 (95% CI 0.63–1.17) — the interval crosses 1, meaning the closest thing this study has to a quasi-experimental cut is not significant.

RCT misses Penda HealthAI Consult / GPT-4oNature Medicine08/01

16 Kenyan clinics, 103 clinical officers, 9,691 patients: GPT-4o as a clinical safety net is "safe but not effective"

What

KEMRI-Wellcome Trust, the University of Birmingham, Penda Health and LSHTM ran a pragmatic cluster-randomized trial across 16 primary care facilities in Nairobi and Kiambu counties, randomizing clinical officers to an electronic medical record with LLM assistance or the same record with the LLM switched off. The tool was AI Consult 2.0, built on GPT-4o. Enrollment ran April 22 – July 16, 2025: 9,691 patients overseen by 103 clinical officers (52 in the LLM-assisted arm, 51 in control). The primary endpoint was an expert-adjudicated composite of treatment failure within 14 days: 102/4,693 (2.2%) in the intervention arm versus 94/4,654 (2.0%) in control, aOR 0.77 (95% CI 0.55–1.08, P = 0.13) — no significant difference. Nature Medicine published it online on June 26 and in the August print issue.

Why it matters

This is the largest and cleanest randomized test to date of putting a frontier LLM straight into a primary care consulting room, and the answer is: it does not hurt anyone, and it does not save anyone either. The authors' own wording: "LLM assistance was safe but did not reduce treatment failure within 14 days and any benefit, if present, is probably modest." Yet every secondary endpoint improved — documentation quality rose significantly (diagnosis aOR 1.74, comprehensive notes aOR 1.68, treatment plans aOR 1.71), mean antibiotic cost fell by US$0.15 per patient, and patient satisfaction was identical. Put plainly: what an LLM can currently prove is documentation and prescribing discipline, not clinical outcomes.

Discount this

Event rates came in far below the design assumption, so the trial is effectively underpowered: secondary write-ups note that detecting a more modest difference would need 100,000-plus patients. "No significant difference" therefore does not mean "proven ineffective." The sample also comes entirely from a single urban private clinic network (Penda Health), limiting generalizability to rural and public systems, and the 14-day follow-up window may be too short. On safety, only 49.4% of flagged outputs were rated "definitely safe" and 1.1% were judged unsafe — the denominator and severity behind that 1.1% need checking against the paper itself.

Quarterly tally 2026 Q3ESTOP-AKIProf. Valmed10/02

A quarterly review published on 10/2 tallies Q3 trial by trial and lands on five words: the workflow is now the model

What

The AI & Digital Health 2026 Q3 review, published October 2, lines up the quarter's clinical evidence, and it reads consistently. ESTOP-AKI (JAMA Network Open, July): in 180 hospitalized patients, machine learning flagged acute kidney injury risk successfully, but earlier nephrology consultation did not improve outcomes (peak creatinine change 0.04 versus −0.03 mg/dL, not significant), and only 41% of non-diet, non-electrolyte recommendations were fully followed. ALLIANCE (September, 82 physicians assessing rheumatology cases) found LLM assistance made clinicians much faster without moving accuracy; ACDIS's coverage supplies the numbers — top-diagnosis accuracy 33.3% with Prof. Valmed versus 35.0% in control, but 94 seconds versus 206 seconds to reach a diagnosis. The same review lists a September hallucination study: junior clinicians caught only 15.8% of GPT-4o hallucinations, and 13.1% of them caught none at all.

Why it matters

Taken together, Q3's shape is clear: models find the risk, but people and processes fail to convert it into action. ESTOP-AKI's 41% adherence and the EDI study's 46.6% notification rate are the same finding wearing different clothes. The review's own verdict is the blunt one: "A better model does not necessarily produce better medicine." That also explains why the quarter's foundation models kept scaling (NeuroVFM on 5.24 million MRI and CT volumes, PRISM2 on 2.3 million whole-slide images and 700,000 pathology reports) with no matching outcome curve beside them.

Discount this

This is a bylined curated review, not a peer-reviewed systematic one: it declares no search strategy and no inclusion or exclusion criteria, so selection bias is possible. Every primary study in it should be checked against its own paper; this issue uses only figures for which a separate secondary or primary source could be found, and flags those that could not be corroborated (such as the full citation for the hallucination study) in the editor's note.

Evidence quality ARISEStanford / Harvard01/15

The number in the Stanford–Harvard report still stands unrefuted: of 500-plus medical AI studies, only 5% used real patient data

What

ARISE (AI Research and Science Evaluation, a Stanford–Harvard research network) published its first annual State of Clinical AI Report 2026, authored by Peter Brodeur, Ethan Goh, Adam Rodman and Jonathan H. Chen, backed by Stanford Computational Medicine, Harvard Medical School's Shapiro Institute and Beth Israel Deaconess. Stanford Medicine's write-up carries the load-bearing figures: the FDA has cleared more than 1,200 AI-enabled medical tools; and of the 500-plus medical AI studies reviewed, nearly 50% used exam-style questions and only 5% used real patient data. The report also records AI predicting patient deterioration 8–24 hours ahead of standard hospital alerts.

Why it matters

That 5% is the denominator for this entire issue. When ninety-five percent of the literature runs on question banks, there is no way to know how much of "progress in medical AI" converts into change at the bedside — and this issue's first three stories happen to show you what that 5% looks like: one with a mortality signal, two without. The report's own line: "AI is already embedded in health care, and that is unlikely to change. What this report makes clear is that the next phase will not be driven by newer models alone." It also names a "jagged frontier": strong on controlled tasks, brittle the moment real-world variability and uncertainty arrive.

Discount this

The report itself was published on January 15 and is not this week's news; it appears here because it is the ruler needed to read the first three stories. ARISE's public page offers only conceptual conclusions, with no full methodology and no list of the 500-plus studies, so the "nearly 50%" and "only 5%" figures can currently be corroborated only through Stanford Medicine's release, not re-checked study by study from the public page.

Cardiac ultrasound Mayo ClinicUltraSightPhilips Lumify06/09

Nine complete novices scanned 995 adults with handheld ultrasound: 90.1% sensitivity, 99.1% specificity — but the two-step pathway cut sensitivity to 77.0%

What

Méabh M. Killalea, Jordan Borgeson, Jared G. Bird and colleagues at Mayo Clinic Rochester published a prospective study in npj Digital Medicine (June 9, 2026; vol 9, article 712): 995 adults underwent both AI-enhanced ECG (AI-ECG) and AI-guided focused cardiac ultrasound (AI-FoCUS), against comprehensive echocardiography as reference. The scans were performed by 9 operators with no prior ultrasound experience, using a Philips Lumify handheld probe with real-time guidance from UltraSight, with blinded expert interpretation. Results: AI-FoCUS detected structural heart disease (SHD) with 90.1% sensitivity, 99.1% specificity, 95.4% PPV and 98.0% NPV; AI-ECG alone reached 87.0% sensitivity but only 60.7% specificity and 30.6% PPV.

Why it matters

This is the issue's only evidence of AI genuinely devolving a skill to non-specialists — and devolving it cleanly, with guidance handling acquisition and experts handling interpretation, so the line of responsibility stays legible. But the figure worth copying down is the trade-off: screening first with AI-ECG and confirming with AI-FoCUS cut imaging utilization to 47% and raised specificity to 99.4% with 96.1% PPV, at the cost of sensitivity falling from 90.1% to 77.0%. Halve the studies, miss roughly a quarter of the patients. That exchange rate — how much saving buys how much missing — is the question payers actually ask.

Discount this

This is a referral cohort with disease prevalence far above community screening, so the 95.4% PPV cannot be carried over to a general population; the authors themselves claim only that it is "feasible and accurate for detecting selected SHD phenotypes" — and "selected" bounds which phenotypes. Note too that novices acquired the images but blinded experts read them, so the study never tests the setting where no expert is available.

Cath lab Powerful MedicalPMcardio / Queen of HeartsJACC: CV Interventions

One AI-ECG took false cath lab activations from 41.8% to 7.9%: a three-network registry of 1,032 STEMI activations

What

The Queen of Hearts US Registry covers three primary PCI networks — UC Davis Sacramento, UT Health Houston and Beth Israel Deaconess Boston — with 1,032 STEMI activations from January 2020 to May 2024. TCTMD's report (the study appeared concurrently in JACC: Cardiovascular Interventions) gives: 92% sensitivity for the AI versus 71% for standard interpretation, 81% specificity versus 29%, AUC 0.94; false-positive cath lab activations among biomarker-negative patients fell from 41.8% to 7.9%. The model is Powerful Medical's PMcardio "Queen of Hearts", trained on occlusion myocardial infarction (OMI) rather than conventional STEMI criteria. Timothy Henry's line: "AI-EKG analysis has the potential to reduce false activation by up to fourfold."

Why it matters

A jump from 29% to 81% specificity is rare in the clinical AI literature, and it points at an underrated fact: much of AI's value lies not in catching more patients but in stopping unnecessary actions. Every false-positive cath lab activation means a team woken at night, a lab tied up, and a patient pushed into an invasive pathway — costs that never appear inside an AUC. Co-investigator Ivan Rokos's line is worth keeping too: "We as physicians are still ultimately responsible for the patient, but AI can certainly turbocharge and augment our abilities."

Discount this

This is a retrospective registry, not a prospective comparison: the AI re-read 1,032 cases that had already been activated, so it never tests what the AI would miss as a first-line gatekeeper. Both the coverage and the original presentation date to October 2025 (TCT 2025), not this week; it appears here to contrast with the previous story, where AI-ECG used alone returned a PPV of just 30.6% — same modality, different decision point, numbers that far apart. Powerful Medical is also the model's commercial developer, so the registry's data provenance and funding relationships need checking against the paper.

Clearance mix FDACureusKaiser Permanente07/13

Of 1,430 FDA AI device clearances over thirty years, radiology took 1,094 (76.5%); pathology got 9, psychiatry zero

What

Pouyan Golshani and colleagues at Kaiser Permanente Los Angeles Medical Center audited FDA AI/ML device clearances from 1995 to 2025 in Cureus (July 13, 2026); AuntMinnie's summary gives the counts: 1,430 devices in total, 1,094 of them radiology (76.5%), with cardiovascular and neurology bringing the top three to 90.6%. Pathology has 9, microbiology 6, obstetrics and gynecology 4, and psychiatry and behavioral health zero. The pace is just as lopsided: a mean of 1.8 clearances a year in 1995–2014 against 264 a year in 2023–2025, with 331 in 2025 alone. Among manufacturers, 502 of 740 companies (67.8%) hold exactly one clearance, while 13 companies (1.8%) hold 247 of them (17.3%).

Why it matters

This distribution explains why all of today's evidence grows on imaging and the heart: in thirty years the regulatory pathway has only been paved for specialties with pixels, signals and a gold standard. Psychiatry's zero is not for want of effort but for want of a reference standard a 510(k) can accept — which is also why most mental health AI ships as a consumer product outside device regulation, bypassing premarket review altogether. Radiology's 76.5% carries a second implication: if you build non-imaging clinical AI, your problem is not only technical, it is that the road has not been walked yet.

Discount this

This is a bibliometric study published July 13, not this week's news — and clearance counts are no proxy for clinical value: the overwhelming majority of the 1,430 are 510(k) substantial-equivalence findings requiring no prospective trial. This issue read only AuntMinnie's secondary summary, which does not break out the cardiovascular and neurology counts, so those are not cited. Separately, the Q3 quarterly review states the FDA has now authorized "more than 1,600" AI-enabled devices; that window differs from the 1,430 here (which stops at end-2025), and the two figures should not be mixed.

Physician attitudes AMA2026 Physician Survey03/12

81% of physicians now use AI themselves, yet 46% do not want patients using AI to read their own radiology results, and 88% fear skill loss

What

The American Medical Association's Center for Digital Health and AI ran its 2026 Physician Survey on Augmented Intelligence, summarized by AuntMinnie (March 12, 2026): 81% of physicians now use AI in practice, up from 38% in 2023, and more than 75% say AI improves their ability to care for patients, up from 65%. Ask about the patient side and it inverts: 46% would "never or rarely" want patients using AI to interpret radiology results, and 49% say the same about pathology results. Separately, 88% worry about skill loss, 86% call data privacy critical to broader adoption, and 88% say the same of safety and efficacy validation.

Why it matters

Set 81% beside 46% and the next phase's real battleground appears: not the model, but who is allowed to point an AI at their own scan. Physicians accept AI as their own copilot and decline it as the patient's — while patient-facing interpretation tools spread fast as consumer products (the same pathway as the previous story's zero psychiatry clearances). That gap will become a concrete clinic-room conflict within a year or two, not an academic argument. And the 88% who fear skill loss line up with today's Prof. Valmed result — 94 seconds versus 206, with accuracy unchanged: faster is not the same as more capable.

Discount this

This is survey coverage from March 12, not this week's news. More importantly, AuntMinnie's report does not disclose the sample size or sampling method, and an AMA membership survey carries self-selection bias, so a figure like "81%" should be read as a self-reported share among respondents, not a population estimate for US physicians. The report also never defines the scope of "using AI" — whether ambient documentation, scheduling and billing count.

02 — Product Analysis

Today deliberately dissects two things that look nothing like rivals: a gradient-boosted score that has existed since 2018, and a 2025 frontier LLM — because the one with a patient outcome is the former

Epic Deterioration Index (EDI)

In-EHR inpatient deterioration risk score · Epic Systems (USA)

Function and position. The EDI is an inpatient deterioration risk score that lives inside the Epic EHR, fed by structured vitals, labs and nursing assessments, emitting a 0–100 score. It is not generative and contains no LLM. Who buys it? Nobody — it is something Epic customers already have, which is why RWJBarnabas's intervention cost sat almost entirely in process redesign rather than procurement. The actual product innovation was changing the score's consumer from a nursing-station dashboard to the rapid response team's pager.

  • Strength : one of the year's very few clinical AI results attached to mortality — 11 hospitals, 23,132 encounters, high-risk inpatient mortality 23.1% → 18.6%, adjusted aOR 0.82 (95% CI 0.74–0.91) — and ICU transfers did not rise proportionally, so it was not bought by shipping everyone to intensive care.
  • Strength : zero procurement friction. It already sits inside thousands of Epic customers' systems, making it far more replicable than any new model requiring contracting, integration and validation — exactly the kind of thing the ARISE report means when it says the next phase will not be driven by newer models alone.
  • Concern : only 46.6% of post-implementation encounters actually fired a notification, so we do not know whether this system becomes an alarm-fatigue disaster at full coverage; and in the band just across the threshold (EDI 55–59) the aOR is 0.86 with a 95% CI of 0.63–1.17 — not significant.
  • Concern : a before-and-after design cannot exclude contributions from concurrent quality-improvement work; two hospitals implemented incompletely, seasons were unevenly distributed, and the authors flag generalizability as uncertain. The missing datum is unambiguous: there is no randomized trial.

AI Consult 2.0 (Penda Health, on GPT-4o)

In-EHR clinical safety net for primary care · Penda Health (Kenya) / OpenAI GPT-4o

Function and position. AI Consult is not a chatbot but a safety-net layer that hooks onto the moment a clinical officer records a diagnosis or writes a prescription and flags likely errors in real time. Its users are primary care clinical officers — non-physician clinicians in Kenya's system — and it runs on GPT-4o. It was designed not to interrupt the workflow, speaking only when it detects a problem; that single design decision ends up explaining both its success and its failure, per the Nature Medicine trial report.

  • Strength : one of very few LLM clinical products willing to run a cluster-randomized trial and publish the result honestly, missed primary endpoint included. That alone makes it more credible than the great majority of medical models that grade their own homework — compare the self-tested, self-scored 100% MedQA problem covered in this site's October 1 issue.
  • Strength : the secondary endpoints genuinely moved — diagnosis documentation aOR 1.74, comprehensive notes aOR 1.68, treatment plans aOR 1.71, antibiotic cost down US$0.15 per patient, with no drop in patient satisfaction. With antimicrobial resistance a global problem, prescribing discipline is itself a reimbursable value.
  • Concern : the primary endpoint missed, and the authors' phrasing caps the ceiling in advance — "any benefit, if present, is probably modest." Worse, it is underpowered: secondary write-ups note that detecting a smaller difference would need 100,000-plus patients, which means products like this cannot produce evidence of reduced treatment failure any time soon — not because the effect is absent, but because it is unshowable at feasible scale.
  • Concern : only 49.4% of flagged outputs were rated "definitely safe" and 1.1% were judged unsafe. Multiply that by another of today's findings — junior clinicians detected only 15.8% of GPT-4o hallucinations — and the assumption that a human will catch the wrong ones does not hold in a primary care setting.

The contrast between the two is this issue's argument. The EDI is far weaker than GPT-4o as a model, yet it attaches to mortality, because its output is wired to a team that physically moves. AI Consult is far stronger, yet it only moved documentation, because its output terminates at the eyes of a busy human being. Clinical AI's bottleneck is no longer the model's ceiling but whether anyone walks the stretch of road after the output.

03 — Companies & Competition

Who stands where, on what, against whom
Company Recent state & numbers Position & moat
Epic Systems
The de facto EHR standard in US hospitals
Its built-in EDI is associated, across RWJBarnabas's 11 hospitals and 23,132 encounters, with high-risk inpatient mortality falling 23.1% → 18.6%, adjusted aOR 0.82, NEJM AI DOI 10.1056/AIoa2500973; RRT evaluation rate 25.3% → 37.5%. Its moat is already being inside: no procurement, no integration, and the distribution rights for clinical AI in its hands. Its weakness is that the model itself is unremarkable — and once the evidence bar rises from association to randomization, Epic has neither the habit nor the incentive to run trials.
Penda Health + OpenAI
Kenyan primary care network × frontier model
A cluster-randomized trial across 16 clinics, 103 clinical officers and 9,691 patients: primary endpoint 2.2% versus 2.0%, aOR 0.77, P = 0.13; documentation quality aOR 1.68–1.74; antibiotic cost down US$0.15 per patient. Its moat is the honesty of its evidence plus real deployment experience in a low-resource setting — OpenAI's hardest-to-copy asset on the clinical side. Its weakness is the business model: gains in documentation quality and antibiotic cost are hard to convert, in Kenyan primary care, into willingness to pay enough to cover GPT-4o inference.
UltraSight
Real-time AI guidance for cardiac ultrasound acquisition
In Mayo's prospective study of 995 adults, 9 operators with zero experience acquired images on a Philips Lumify, and AI-FoCUS detected SHD with 90.1% sensitivity, 99.1% specificity and 95.4% PPV (interpretation by blinded experts). Its moat is commoditizing the most hand-skill-dependent step in echocardiography, bundled into Philips's handheld hardware channel. Its weakness is that it solves acquisition and not interpretation, so what it sells is a preprocessor for experts rather than a substitute — and in regions short of cardiologists, including parts of rural Taiwan, that limit is its ceiling.
Powerful Medical
PMcardio / Queen of Hearts, OMI-oriented AI-ECG
A retrospective registry of 1,032 STEMI activations across three networks: 92% sensitivity versus 71%, 81% specificity versus 29%, AUC 0.94, with false-positive cath lab activations among biomarker-negative patients falling 41.8% → 7.9%. Its moat is choosing a different training target — occlusion MI rather than conventional STEMI criteria — so it catches the group textbook criteria miss. Its weakness is the retrospective design, plus the fact that what it mainly improves is avoiding unnecessary activations: a value that is harder to sell in systems paid by procedure volume.
ARISE(Stanford × Harvard)
Not selling a product — selling the evaluation standard
Its first annual report notes the FDA has cleared more than 1,200 AI-enabled devices while, among the 500-plus studies reviewed, nearly 50% used exam questions and only 5% used real patient data; it argues the evidence bar must move to multi-turn unstructured data and real-world consequences. Its moat is academic neutrality and joint endorsement from Stanford, Harvard and BIDMC — in a market where vendors routinely grade themselves, defining what counts as effective is itself a form of power. Its weakness is that it has no enforcement, and no published full methodology to re-check study by study.

Today's competitive structure is a crossing of distribution against evidence. Whoever holds distribution (Epic) can deploy without evidence, while whoever produces the most honest evidence (Penda and OpenAI) holds no distribution. The layer in between — UltraSight, Powerful Medical — has each chosen to solve one short stretch of workflow and leave responsibility with the expert: regulatorily safe, and a ceiling welded on at the same time.

04 — Taiwan Angle

Taiwan is building data infrastructure — and today's evidence says what will be missing once it is built is workflow

(1) The timeline is now pinned to year-end — but it is a data timeline, not an outcomes timeline. Health Minister Shih Chung-liang, previewing the Biotech Forum held on October 7, said on September 15 that medical AI's "first task is establishing data standards, governance systems and a trustworthy operating environment," and gave firm dates: FHIR data readiness at medical centers by the end of 2026, and nationwide hospital interconnection by the end of 2027. The policy instruments include the five-year, NT$489 billion Deepening Health Taiwan Plan, the four-year, NT$240 billion National Drug Resilience Preparedness Plan, and a planned NT$100 billion investment in innovative medical devices and drug R&D (Economic Daily News / UDN, 2026-09-15). Set that against today's RWJBarnabas result: they had no new data and no new model — they wired an existing score to the rapid response team and got aOR 0.82. FHIR interconnection is a necessary condition, not a sufficient one.

(2) Taiwan already has the institutional skeleton of "responsible AI centers" — what it lacks is the RRT wire. In 2025 the Ministry of Health and Welfare established responsible-AI implementation centers with ten hospitals, framed as "trustworthy gatekeepers for smart healthcare" (MOHW). The gatekeeper metaphor is itself the problem: gatekeeping asks whether a model may enter, while today's evidence points entirely at what happens afterwards — who receives the alert, and who moves. Only 41% of ESTOP-AKI's recommendations were fully followed; the EDI fired in only 46.6% of encounters. If Taiwan's medical centers treat AI adoption as an IT department project rather than a quality-improvement project, that same gap copies over unchanged.

(3) On the payment side, what is most worth copying today is the studies avoided, not the patients caught. Mayo's two-step pathway cut imaging utilization to 47% at the cost of sensitivity falling from 90.1% to 77.0% (npj Digital Medicine); Powerful Medical cut false-positive cath lab activations from 41.8% to 7.9% (TCTMD). Under Taiwan's global-budget NHI, "fewer unnecessary procedures" converts into a reimbursement argument more easily than "higher detection rate" — and is also more easily misused as a pretext for squeezing test volumes. The sensitivity that fell from 90.1% to 77.0% is precisely the cost that has to be stated out loud before that ledger is opened.

05 — Further Reading

Chosen on one criterion: it changes the question you ask of the next AI paper, rather than handing you one more number
  1. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial — Nature Medicine (2026-06-26)

    Copy the abstract's closing line and use it as a ruler: "LLM assistance was safe but did not reduce treatment failure within 14 days." This is the cleanest randomized evidence yet for a frontier model in a real consulting room, and the methods deserve a full read — above all the passage on event rates coming in below expectation and leaving the trial underpowered.

  2. State of Clinical AI Report 2026 — ARISE, Stanford–Harvard Research Network

    Read it not for the 5% but for the "jagged frontier" frame: a single model's gap between controlled tasks and real-world variability is not a smooth slope but a sawtooth. That frame changes how you read any benchmark.

  3. AI & Digital Health: Quarterly Review 2026Q3 — micheledpierri.com (2026-10-02)

    The freshest document in this issue, and the only one that tallies Q3's trials, foundation models and regulatory moves side by side. Note that it files ESTOP-AKI's 41% adherence and ALLIANCE's faster-but-no-more-accurate result in the same section — that arrangement is itself the argument.

  4. Advancing conversational diagnostic AI with multimodal reasoning — Nature Medicine (2026-05-14)

    Multimodal AMIE beat primary care physicians on 29 of 32 evaluation axes across 105 simulated telehealth consultations — and then the authors themselves note it is not a clinical trial, that a chat interface has no physical examination, and that pretraining data leakage is possible. Read it against this issue's second story and the chasm between winning an OSCE and not winning an RCT comes into focus.

  5. An agentic system for rare disease diagnosis with traceable reasoning (DeepRare) — Nature, vol 651, 775–784 (2026)

    Rare disease is one of the few settings where AI should genuinely hold an advantage: prevalence under 1 in 2,000, patterns clinicians were never going to connect, and diagnostic odysseys routinely running five years or more. Read it for the traceable-reasoning design — tying the inference chain back to medical evidence is the only approach in this issue that confronts head-on how a human is supposed to check an AI.

06 — References

References
  1. Artificial intelligence (AI)-triggered rapid response notifications were associated with lower mortality among inpatients. 2 Minute Medicine, 2026-10-02. 2minutemedicine.com
  2. RWJBarnabas Links Real-Time Epic EDI Alerts Directly to Rapid Response Teams (NEJM AI, DOI 10.1056/AIoa2500973). HIT Consultant, 2026-07-29. hitconsultant.net
  3. AI-enabled tool helps identify hospitalized patients at risk of rapid clinical decline. News-Medical, 2026-07-29. news-medical.net
  4. RWJBarnabas Health Sees Improvements Using Epic Deterioration Index. Healthcare Innovation, 2026-07. hcinnovationgroup.com
  5. AI tool helps detect patient deterioration earlier, reducing hospital deaths. Medical Xpress, 2026-07. medicalxpress.com
  6. Agweyu A, Mwaniki P, Menon V, et al. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine, online 2026-06-26, print 2026-08. nature.com
  7. PubMed record 42362867 (full abstract, Nature Medicine). PubMed / Europe PMC, 2026. pubmed.ncbi.nlm.nih.gov
  8. AI clinical support tool improved clinician decisions in real-world primary care trial. University of Birmingham, 2026. birmingham.ac.uk
  9. Generative AI Clinical Decision Support Shows Safety and Documentation Gains But No Significant Reduction in Treatment Failure in Kenyan Primary Care Trial. MedPath, 2026. trial.medpath.com
  10. AI & Digital Health: Quarterly Review 2026Q3. micheledpierri.com, 2026-10-02. micheledpierri.com
  11. News: AI tool quicker, not more accurate than humans when diagnosing rheum disease, study finds (medRxiv, 2026-08-29). ACDIS, 2026-09-10. acdis.org
  12. Brodeur P, Goh E, Rodman A, Chen JH. State of Clinical AI Report 2026. ARISE (Stanford–Harvard Research Network), 2026. arise-ai.org
  13. Clinical AI Has Boomed. A New Stanford-Harvard State of Clinical AI Report Shows What Holds Up in Practice. Stanford Medicine, 2026-01-15. medicine.stanford.edu
  14. Killalea MM, Borgeson J, Bird JG, et al. Improving structural heart disease screening: AI-ECG and novice AI-guided focused cardiac ultrasound. npj Digital Medicine 9:712, 2026-06-09. nature.com
  15. AI-ECG Finds STEMI Faster, Cuts False-Positive Cath Lab Activations (Queen of Hearts US Registry; JACC: Cardiovascular Interventions). TCTMD, 2025-10-31. tctmd.com
  16. Queen of Hearts: AI ECG for STEMI & OMI. Powerful Medical. powerfulmedical.com
  17. Radiology dominates thirty years of FDA AI device approvals (Golshani P, et al. Cureus, 2026-07-13). AuntMinnie, 2026-07-16. auntminnie.com
  18. Physicians mixed on patient use of AI to interpret radiology results (AMA 2026 Physician Survey on Augmented Intelligence). AuntMinnie, 2026-03-12. auntminnie.com
  19. Advancing conversational diagnostic AI with multimodal reasoning. Nature Medicine, 2026-05-14. nature.com
  20. Zhao et al. An agentic system for rare disease diagnosis with traceable reasoning (DeepRare). Nature 651:775–784, 2026. nature.com
  21. AI succeeds in diagnosing rare diseases (News & Views). Nature, 2026-02-18. nature.com
  22. 生技論壇10月7日登場 衛福部長石崇良:積極建構AI醫療基建. 經濟日報/聯合新聞網, 2026-09-15. udn.com
  23. 醫療AI要聰明,還要能負責!十家醫院打造可信賴的智慧醫療守門人. 衛生福利部, 2025. mohw.gov.tw
  24. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine, 2026. nature.com
  25. Registered AI Medical-Imaging Clinical Trials on ClinicalTrials.gov: Publication Yield, Predictors, and Portfolio Evolution. Academic Radiology, 2026 (abstract not retrievable; see editor's note). academicradiology.org
Editor's note: (1) News on this rotation was thin today — no new large prospective clinical AI trial was published in the past 72 hours, so this issue extends the window back to September 28 and builds its spine from two documents published on October 2 (2 Minute Medicine's appraisal of the NEJM AI study, and the Q3 quarterly review). Every other story carries its original publication date in its tags and source row; the oldest is TCTMD's October 2025 report, kept deliberately to contrast one modality's performance at two different decision points. (2) Pages that could not be read: the Nature Medicine article page (s41591-026-04503-6) and a Europe PMC bibliographic query for the Academic Radiology paper were refused by the proxy with HTTP 429 during this run, while PubMed and academicradiology.org returned reCAPTCHA and 403 respectively. The Kenyan trial's full abstract and statistics come from the Europe PMC record for PubMed 42362867 (verbatim abstract); its secondary endpoints (aOR 1.68–1.74, US$0.15, 49.4% / 1.1%) are cited from MedPath's secondary summary rather than the paper, so defer to the original. For the Academic Radiology paper only the title was obtained and no figures are cited. The Lancet Digital Health's editorial (PIIS2589-7500(26)00053-1) also returned 403 and was therefore left out of Further Reading. (3) Vendor-reported figures and design limits: Powerful Medical is Queen of Hearts' commercial developer and that registry is retrospective; UltraSight and Philips supplied the equipment in the AI-FoCUS study. The EDI study is a before-and-after cohort, not a randomized trial. (4) Proportions that could not be re-checked study by study: ARISE's "nearly 50% exam questions / only 5% real patient data" is corroborated only by Stanford Medicine's release, as the public page lists no methodology; the Q3 review is a bylined curation rather than peer-reviewed, and original citations could not be obtained for its September hallucination study (15.8% / 13.1%) or for some ESTOP-AKI and ALLIANCE details — ALLIANCE's accuracy and timing figures are separately corroborated by ACDIS. (5) The AMA survey's sample size and sampling method are not disclosed in the secondary coverage. (6) This report is not investment or medical advice.