◆ AI & Medical AI Daily
–
Monday · Clinical Applications & Research

AI helps — but not always the person you expected

The most interesting turn in clinical AI research this year is that the question has shifted from "is the AI accurate?" to "who does it actually help, and how must it be presented?" A 1,096-participant randomised study in Nature Medicine on Aug 4 gives the bluntest answer yet: the same dermatology AI moved lay people and primary care physicians in opposite directions — LLM narrative explanations gained lay users 13.4% when the AI was right and cost them 21.1% when it was wrong, while for physicians they were the least useful explanation format yet the only one that improved confidence calibration. The same week, Aidoc's CARE Body CT Multi-Triage became the first healthcare-foundation-model AI software to win a Medicare NTAP add-on payment, effective Oct 1. Evidence and money are being pushed forward by two very different logics.

01 — Top Stories

Seven things in clinical AI
Research Nature MedicineMITDermatology8/4

One dermatology AI, two opposite outcomes: a 1,096-person randomised study takes apart explainable AI

What

Marzyeh Ghassemi's group at MIT reports two large randomised experiments in Nature Medicine vol. 32 issue 8 (published online Aug 4, 2026, open access), using a 4×2 factorial design: four explanation formats (bare prediction, GradCAM heatmap, similar-case retrieval, LLM narrative) crossed with two decision orders (human-first vs AI-first). Participants were 623 lay people on binary melanoma classification and 153 primary care physicians on complex differential diagnosis, with 320 medical students as a comparison group.

Why it matters

The findings read like a purpose-built warning about the current trend of patients pointing AI at their own skin photos. For lay users: AI lifted accuracy from 69.7% to 75.8%, but LLM narratives cut both ways — +13.4% when the AI was correct, −21.1% when it was not, amplifying automation bias. For physicians: top-1 accuracy rose 21.5% and the benefit persisted even under incorrect suggestions; LLM explanations helped least on accuracy (+17.7%) yet were the only format that improved confidence calibration. Both groups shared good news — diagnostic disparities across skin tones narrowed (46.9% relative reduction for lay users, 35.6% for physicians) — and bad news: showing the AI answer first significantly increased deference and anchoring. The ordering of a user interface, in other words, matters on the same order of magnitude as the model's accuracy.

Players

Corresponding author Marzyeh Ghassemi (MIT), co-first authors Xuhai "Orson" Xu and Haoyu Hu, with more than 20 co-authors (full text in Nature Medicine). No commercial product is under test — the study evaluates four generic explainability techniques.

RCT Retina4IRDNature MedicineOphthalmology

A multicentre RCT in inherited retinal disease: clinician genetic-diagnosis accuracy goes 67.3% → 88.5%

What

A multicentre randomised trial led by Xiaodong Sun's group at Shanghai General Hospital (Shanghai Jiao Tong University), spanning China, South Korea and Poland, published in Nature Medicine on July 24, 2026. The development cohort comprised 1,843 genetically confirmed patients (3,376 eyes); the randomised cohort was 300 people with suspected inherited retinal disease (IRD), allocated 1:1, with 295 analysed (median age 33, 38.6% female). Retina4IRD integrates multimodal retinal imaging and reached 90.4% top-5 accuracy on internal validation and 85.6% externally.

Why it matters

IRD is the archetypal "rare but increasingly treatable" disease family — gene therapies and trial slots keep expanding, yet naming the causative gene demands scarce subspecialty experience. What matters here is not only accuracy: clinicians in the AI-assisted arm reached 88.5% genetic-diagnosis accuracy versus 67.3% in the control arm — a 21.2 percentage-point gap, p < 0.001 — with top-1 accuracy of 37.8% vs 22.4%. More importantly, composite downstream management scores also improved significantly (37.7 vs 28.5, p < 0.001). Most medical-AI trials stop at read accuracy; this one measures what the clinician then did, which is methodologically unusual.

Players

The team spans Shanghai Jiao Tong University, Nanjing University of Aeronautics and Astronautics, Tsinghua University, Yonsei University and several Polish institutions (full text). Retina4IRD is an academic system; no public commercialisation or regulatory filing was found, and none is assumed here.

Reimbursement AidocCMSNTAP8/13

Aidoc's CARE becomes the first healthcare-foundation-model AI software to win Medicare NTAP, effective Oct 1

What

CMS granted a New Technology Add-on Payment to Aidoc's CARE (Clinical AI Reasoning Engine) Body CT Multi-Triage, reported by AuntMinnie on Aug 13, effective Oct 1, 2026 for three years, covering qualifying Medicare fee-for-service inpatient cases. The device flags suspected urgent findings on chest, abdomen and pelvis CT (with and without contrast) for workflow triage; it received FDA Breakthrough Device Designation in September 2025 and FDA clearance in January 2026 (Aidoc release, Radiology Business, Diagnostic Imaging).

Why it matters

AuntMinnie frames it plainly as "a novel reimbursement decision for AI software built on a healthcare foundation model." NTAP has historically gone to narrow, single-indication algorithms; awarding it to a product that reads multiple urgent conditions off one scan means CMS has effectively recognised multi-task AI as a payable category. The practical clinical consequence: imaging AI shifts from a hospital-funded efficiency bet to a line item with cash flow behind it, sharply lowering the financial barrier to adoption. It also runs straight into the concern researchers raised in STAT's same-day piece: a per-use top-up rewards volume, and when it expires in two to three years the product must justify itself on clinical benefit alone.

Players

Aidoc (Tel Aviv, Israel) is among the market leaders in imaging triage AI and in January 2026 expanded its CT triage platform to 14 indications, 11 of them new. CARE is its foundation-model line. CMS has not published the per-case add-on amount and none is inferred here (original release).

Frontline National Nurses UnitedKaiserMontefiore8/11

Nurses want a seat at the clinical-AI table — from strikes to co-design

What

STAT reported on Aug 11 (by Katie Palmer) that US nursing's pushback on clinical AI is becoming institutional. Two flashpoints are named: laid-off nurses at Montefiore in the Bronx questioning administrative AI displacing staff, and nurses at Kaiser Permanente in California striking over AI surveillance of their work. National Nurses United, representing more than 200,000 nurses, is the main actor.

Why it matters

Clinical AI increasingly fails not because the model is inaccurate but because nobody uses it as designed after deployment. Nurses are the frontline users of bedside alerts, deterioration detection and documentation AI — and among the least studied. Story 7 below makes the point concretely: in the UVA deterioration trial, clinicians moved 11% of patients between display and non-display beds, letting clinical judgement override the randomisation outright. STAT's argument is to convert an adversarial relationship into co-design — which is why "who was involved in deployment" is becoming as consequential a variable as model performance.

Sources: STAT · Aug 11 (STAT+ paywall; this item is written only from the publicly visible headline and lede)
Infrastructure Nature MedicineNordic8/14

Five Nordic countries plus Estonia build a shared AI-Health platform on population registries

What

Ole A. Andreassen of the University of Oslo and 21 co-authors set out the technical foundations, data assets and deployment roadmap of the Nordic AI-Health Initiative in a Nature Medicine Comment (Aug 14, 2026). Participating institutions span Norway, Denmark, Iceland, Sweden, Finland and Estonia, including Karolinska Institutet, the University of Helsinki and Amgen's deCODE genetics. The stated aim is to "enable secure, regulation-compliant access and to deliver generalizable models for responsible AI-driven discovery in medicine."

Why it matters

Nordic longitudinal population registries — cradle-to-grave, multi-generational, linkable — are a globally scarce asset that GDPR and cross-border governance have made hard to use jointly. The significance is not any single model but that this is among the first regional attempts, post-EU AI Act, to build compliance into the data infrastructure itself rather than bolt it on afterwards. The reference value for Taiwan is high: the NHI database has the same longitudinal advantage and the same cross-institutional governance bottleneck (see the Taiwan Angle below). Note this is a Comment, not primary research — no cohort sizes or dated milestones are published.

Benchmark RSNACleveland ClinicKaggle8/5

RSNA opens a $77,000 knee-MRI AI challenge — the first to pair images with reports in 12 languages

What

RSNA announced the 2026 Knee Abnormality Detection AI Challenge on Aug 5, 2026. The dataset holds more than 5,000 knee MRI exams from 16 sites worldwide, with radiology reports in 12 languages. Total prize money is $77,000, including awards for the most efficient models; the challenge runs on Kaggle through Oct 22, winners announced in November and recognised at RSNA 2026 (Nov 29–Dec 3, Chicago). Co-leads are Po-Hao Chen, MD, MBA and Naveen Subhas, MD, MPH of Cleveland Clinic (AuntMinnie).

Why it matters

RSNA calls it "the most real-world challenge yet" precisely because it withholds clean structured labels and forces entrants to learn from messy, variably formatted, multilingual reports. That matters disproportionately for non-anglophone health systems, Taiwan included: most imaging-AI training corpora and benchmarks are English-report-based, and language mismatch is the most commonly underestimated source of performance drop at deployment. The second signal is the efficiency prize — academic benchmarks are finally treating inference cost as a first-class metric rather than optimising AUC alone.

Negative result UVACoMETNihon Kohden

A 10,422-visit deterioration-detection cluster RCT: no difference in the primary outcome — and the randomisation broke

What

Jessica Keim-Malpass, Jamieson M. Bourque and colleagues at the University of Virginia School of Medicine published a cluster-randomised controlled trial in Scientific Reports (Feb 5, 2026) covering 10,422 inpatient visits on an 85-bed cardiology and cardiac surgery ward. The intervention was CoMET (Continuous Monitoring of Event Trajectories) from Nihon Kohden Digital Health Solutions, which passively displays a risk trajectory refreshed every 15 minutes from machine-learning models over cardiorespiratory monitoring and vital signs. The primary outcome was hours free of clinical deterioration (death, emergent ICU transfer, emergent intubation, cardiac arrest, emergent surgery) within 21 days of admission.

Why it matters

"There was no change in the primary outcome between groups" — passively displaying a risk score did not change patient outcomes. Two methodological details deserve more attention than the headline. Only 5.3% of patients had a deterioration event, so power was structurally thin; and clinicians moved 11% of patients between display and non-display beds, preferentially shifting sicker patients into monitored beds, compromising the randomised design outright. This is exactly the bind flagged in Nature's July 28 comment, "Medical AI has a measurement problem" (Arjun K. Manrai, Harvard Medical School): evidence-based medicine assumes the yardstick is reliable, and in the AI setting that assumption is often shaky. Clinical AI that only displays and does not redesign the workflow and escalation path tends to produce precisely this result.

02 — Product Analysis

Two products, taken apart: multi-task triage vs rare-disease decision support

Aidoc CARE Body CT Multi-Triage

Imaging triage · healthcare foundation model · Aidoc (Israel)

Function and position. Detects multiple suspected urgent conditions on chest/abdomen/pelvis CT, with and without contrast, and reorders the worklist — "look at the right scan first," not autonomous interpretation (AuntMinnie).

Retina4IRD

Rare eye disease · multimodal decision support · Shanghai Jiao Tong University (academic)

Function and position. Integrates multimodal retinal imaging to rank candidate causative genes for suspected inherited retinal disease, guiding the clinician's subsequent genetic testing and management (Nature Medicine).

03 — Companies & Competition

Who is selling what, on what evidence, against whom
Company / system Evidence Position
Aidoc
CARE Body CT Multi-Triage
Complete regulatory and payment path (breakthrough Sept 2025, clearance Jan 2026, NTAP Oct 1, 2026); public prospective outcome data is limited. Platform play across indications — 14 CT triage indications; the moat is deployment scale plus first-mover status on reimbursement.
ScreenPoint Medical
Transpara v1.7
The strongest set: a 31,301-woman prospective paired non-inferiority trial (Nature Medicine, Mar 19, 2026) — 63.6% fewer radiologist readings, cancer detection +15.2% (6.3→7.3 per 1,000), recall +14.8% (4.8%→5.5%). Deep in a single indication (breast screening), winning on randomised evidence, with the Swedish MASAI trial's interval-cancer and sensitivity data behind it.
Nihon Kohden
CoMET
Has completed a large cluster RCT, but the result was negative — 10,422 visits, no primary-outcome difference, with 11% of patients moved across arms. A bedside-monitoring hardware vendor extending into predictive analytics: distribution through installed devices is the advantage, passive display is the design weakness.
Scanslated
patient-friendly reports
Large real-world usage data: 391,713 exams at Stanford and Duke children's hospitals (Jan 2024–Nov 2025) — 92.6% of families reported better comprehension, but only 9% opened the patient-friendly version versus 50.5% for the standard report. Attacks the neglected patient-comprehension layer; the hospital pitch is satisfaction and retention (91.8% of families more likely to return) rather than diagnostic performance.
One cross-cutting observation. Evidence strength across these four is strikingly uneven: the company with the best payment terms (Aidoc) is not the one with the strongest public clinical evidence (ScreenPoint), and the one that ran the most complete randomised trial (Nihon Kohden/UVA) got a null result. A Nature Medicine benchmark adds pressure from another direction: general-purpose LLMs outperform specialised clinical AI tools on medical benchmarks, which is a real threat to the long-run moat of narrow, purpose-built models.

04 — Taiwan Angle

Taiwan: three places where today's news lands

(1) The multilingual-report challenge is an opening for Taiwanese datasets. RSNA's knee-MRI challenge is the first to include radiology reports in 12 languages from 16 institutions, an implicit admission that non-English reporting is the real-world norm. Taiwan's mixed Chinese-English reports — English structured fields, Chinese narrative — are essentially absent from international benchmarks, and language mismatch is a common source of the performance drop seen when imported imaging AI is deployed locally. Entering the challenge, or building a local multilingual benchmark, is an actionable move either way.

(2) The Nordic platform's governance design is more worth copying than its models. Taiwan's NHI database shares the Nordic registries' longitudinal scarcity value — and their bottleneck: cross-institutional governance and compliance. The Nordic AI-Health Initiative (Aug 14) writes compliance into the infrastructure rather than bolting it on, which is a ready-made comparator for MOHW's Guidelines on Generative AI Use in Healthcare Institutions and the existing three national smart-healthcare centres, including the Clinical AI Validation Center.

(3) Payment is the real adoption gate — and Taiwan's box is still empty. Aidoc's NTAP means US hospitals now have cash flow behind imaging-AI adoption. Taiwan's NHI still has no corresponding AI-software payment category, so deployment costs sit almost entirely with hospitals. Health Minister Shih Chung-liang set out a "333 policy" and an NT$48.9 billion budget at the Kaohsiung Medical University forum (June 27, 2026), targeting record interoperability across academic medical centres by year-end and regional and district hospitals within two years; the same event highlighted KMU's NVIDIA-based colorectal cancer detection system built with Foxconn. Interoperability is necessary — but without a payment design behind it, the Aidoc pattern of "adequate evidence, viable finances" stays hard to replicate here.

05 — Further Reading

Five worth a second read
  1. Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people — Nature Medicine, 2026/8/4

    The one to read end-to-end. Open access, with the full 4×2 factorial design and every subgroup analysis; anyone designing a clinical AI interface should start with the methods and Figure 3.

  2. Medical AI has a measurement problem — Nature, 2026/7/28

    Arjun K. Manrai's short comment states the core problem most clearly — evidence-based medicine assumes a reliable yardstick, and AI undermines that assumption. It is the frame for this year's run of null trials.

  3. AI-based triage and decision support in mammography and digital tomosynthesis: a paired, noninferiority trial — Nature Medicine, 2026/3/19

    A 31,301-woman prospective trial and the hardest evidence yet that imaging AI cuts reading volume — but recall also rose 14.8%, which belongs in the same sentence as the benefit.

  4. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks — Nature Medicine

    If general models keep winning on benchmarks, what exactly is the moat for narrow clinical AI? Required background for any medical-AI product strategy discussion.

  5. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial — Nature Medicine

    Pragmatic cluster randomisation is the closest thing to a real-world test of clinical AI. Read alongside the UVA CoMET trial to see whether the bottleneck is intervention design or model performance.

06 — References

References
  1. Xu, X., Hu, H., … Ghassemi, M. “Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people.” Nature Medicine, vol. 32 no. 8, 2026-08-04. nature.com/articles/s41591-026-04553-w
  2. Jia, H., Qian, B., Sun, X. “AI-based clinician decision support system for diagnosis of inherited retinal diseases: a multicenter, randomized trial.” Nature Medicine, 2026-07-24. nature.com/articles/s41591-026-04545-w
  3. “CMS approves Medicare add-on payment for Aidoc CT triage AI.” AuntMinnie, 2026-08-13. auntminnie.com
  4. “Aidoc's CARE Body CT Multi-Triage receives eligibility for Medicare New Technology Add-on Payment.” PR Newswire (Aidoc press release). prnewswire.com
  5. “Medicare approves new technology add-on payment for inpatient radiology AI solution.” Radiology Business. radiologybusiness.com · “CT Triage Software Garners NTAP Reimbursement from CMS.” Diagnostic Imaging. diagnosticimaging.com
  6. Palmer, K. “Nurses seek a seat at the table as they fight expanding clinical AI.” STAT, 2026-08-11 (STAT+ paywall). statnews.com
  7. Palmer, K. “What Medicare incentives for AI-based devices mean for tech companies — and hospitals.” STAT, 2026-08-13 (STAT+ paywall). statnews.com
  8. Andreassen, O. A., Heimer, H., Kallioniemi, O., et al. “An AI-Health infrastructure for the Nordic region: technical foundations, data assets, and a roadmap for deployment.” Nature Medicine (Comment), 2026-08-14. nature.com/articles/s41591-026-04575-4
  9. “RSNA Launches Knee Abnormality Detection AI Challenge.” RSNA News, 2026-08-05. rsna.org · challenge page: rsna.org/artificial-intelligence · AuntMinnie coverage: auntminnie.com
  10. Keim-Malpass, J., Bourque, J. M., et al. “A randomized controlled trial of artificial intelligence-based analytics for clinical deterioration.” Scientific Reports, 2026-02-05. nature.com/articles/s41598-026-39051-z
  11. Manrai, A. K. “Medical AI has a measurement problem.” Nature, 2026-07-28. nature.com/articles/d41586-026-02125-z
  12. Álvarez-Benito, M., et al. “AI-based triage and decision support in mammography and digital tomosynthesis for breast cancer screening: a paired, noninferiority trial.” Nature Medicine, 2026-03-19. nature.com/articles/s41591-026-04277-x
  13. “Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading (MASAI).” The Lancet. thelancet.com
  14. “General-purpose large language models outperform specialized clinical AI tools on medical benchmarks.” Nature Medicine. nature.com/articles/s41591-026-04431-5 · “Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial.” Nature Medicine. nature.com/articles/s41591-026-04503-6
  15. Johnston, A., et al. “Patient-friendly radiology reports improve family comprehension” (Pediatric Radiology, reported 2026-08-03). AuntMinnie. auntminnie.com
  16. 「高醫大論壇揭示AI醫療新局!衛福部推『333政策』」,聯合新聞網,2026-06-27。udn.com ·「衛生福利部頒布『醫療機構應用生成式人工智慧指引』」,理律法律事務所。leeandli.com · 臺灣智慧醫療三大中心 aicenter.mohw.gov.tw
  17. “FDA clears CT-based AI triage platform Aidoc.” Diagnostic Imaging, 2026-01-21. diagnosticimaging.com
Editor's note: (1) Aug 15–17 fell on a weekend with very little newly published primary clinical-AI research, so this edition widens coverage to the two weeks from Aug 3 and adds earlier 2026 trials that serve as necessary comparators (the breast-screening non-inferiority trial of Mar 19 and the UVA deterioration trial of Feb 5); all dates are labelled. (2) Both STAT pieces (Aug 11 on nurses, Aug 13 on NTAP) sit behind the STAT+ paywall; this report draws only on the publicly visible headline, standfirst and summary, and cites no figure it could not verify. (3) CMS has not published a per-case NTAP amount for Aidoc CARE, and none is inferred here. (4) Retina4IRD is an academic system; no public commercialisation or regulatory filing was found during verification. (5) Taiwan's "333 policy" and budget figure come from the June 27 United Daily News report; the original states NT$48.9 billion (written 489 億元 in Chinese counting units) — the two renderings are the same number.