AI helps — but not always the person you expected
The most interesting turn in clinical AI research this year is that the question has shifted from "is the AI accurate?" to "who does it actually help, and how must it be presented?" A 1,096-participant randomised study in Nature Medicine on Aug 4 gives the bluntest answer yet: the same dermatology AI moved lay people and primary care physicians in opposite directions — LLM narrative explanations gained lay users 13.4% when the AI was right and cost them 21.1% when it was wrong, while for physicians they were the least useful explanation format yet the only one that improved confidence calibration. The same week, Aidoc's CARE Body CT Multi-Triage became the first healthcare-foundation-model AI software to win a Medicare NTAP add-on payment, effective Oct 1. Evidence and money are being pushed forward by two very different logics.
01 — Top Stories
One dermatology AI, two opposite outcomes: a 1,096-person randomised study takes apart explainable AI
Marzyeh Ghassemi's group at MIT reports two large randomised experiments in Nature Medicine vol. 32 issue 8 (published online Aug 4, 2026, open access), using a 4×2 factorial design: four explanation formats (bare prediction, GradCAM heatmap, similar-case retrieval, LLM narrative) crossed with two decision orders (human-first vs AI-first). Participants were 623 lay people on binary melanoma classification and 153 primary care physicians on complex differential diagnosis, with 320 medical students as a comparison group.
The findings read like a purpose-built warning about the current trend of patients pointing AI at their own skin photos. For lay users: AI lifted accuracy from 69.7% to 75.8%, but LLM narratives cut both ways — +13.4% when the AI was correct, −21.1% when it was not, amplifying automation bias. For physicians: top-1 accuracy rose 21.5% and the benefit persisted even under incorrect suggestions; LLM explanations helped least on accuracy (+17.7%) yet were the only format that improved confidence calibration. Both groups shared good news — diagnostic disparities across skin tones narrowed (46.9% relative reduction for lay users, 35.6% for physicians) — and bad news: showing the AI answer first significantly increased deference and anchoring. The ordering of a user interface, in other words, matters on the same order of magnitude as the model's accuracy.
Corresponding author Marzyeh Ghassemi (MIT), co-first authors Xuhai "Orson" Xu and Haoyu Hu, with more than 20 co-authors (full text in Nature Medicine). No commercial product is under test — the study evaluates four generic explainability techniques.
A multicentre RCT in inherited retinal disease: clinician genetic-diagnosis accuracy goes 67.3% → 88.5%
A multicentre randomised trial led by Xiaodong Sun's group at Shanghai General Hospital (Shanghai Jiao Tong University), spanning China, South Korea and Poland, published in Nature Medicine on July 24, 2026. The development cohort comprised 1,843 genetically confirmed patients (3,376 eyes); the randomised cohort was 300 people with suspected inherited retinal disease (IRD), allocated 1:1, with 295 analysed (median age 33, 38.6% female). Retina4IRD integrates multimodal retinal imaging and reached 90.4% top-5 accuracy on internal validation and 85.6% externally.
IRD is the archetypal "rare but increasingly treatable" disease family — gene therapies and trial slots keep expanding, yet naming the causative gene demands scarce subspecialty experience. What matters here is not only accuracy: clinicians in the AI-assisted arm reached 88.5% genetic-diagnosis accuracy versus 67.3% in the control arm — a 21.2 percentage-point gap, p < 0.001 — with top-1 accuracy of 37.8% vs 22.4%. More importantly, composite downstream management scores also improved significantly (37.7 vs 28.5, p < 0.001). Most medical-AI trials stop at read accuracy; this one measures what the clinician then did, which is methodologically unusual.
The team spans Shanghai Jiao Tong University, Nanjing University of Aeronautics and Astronautics, Tsinghua University, Yonsei University and several Polish institutions (full text). Retina4IRD is an academic system; no public commercialisation or regulatory filing was found, and none is assumed here.
Aidoc's CARE becomes the first healthcare-foundation-model AI software to win Medicare NTAP, effective Oct 1
CMS granted a New Technology Add-on Payment to Aidoc's CARE (Clinical AI Reasoning Engine) Body CT Multi-Triage, reported by AuntMinnie on Aug 13, effective Oct 1, 2026 for three years, covering qualifying Medicare fee-for-service inpatient cases. The device flags suspected urgent findings on chest, abdomen and pelvis CT (with and without contrast) for workflow triage; it received FDA Breakthrough Device Designation in September 2025 and FDA clearance in January 2026 (Aidoc release, Radiology Business, Diagnostic Imaging).
AuntMinnie frames it plainly as "a novel reimbursement decision for AI software built on a healthcare foundation model." NTAP has historically gone to narrow, single-indication algorithms; awarding it to a product that reads multiple urgent conditions off one scan means CMS has effectively recognised multi-task AI as a payable category. The practical clinical consequence: imaging AI shifts from a hospital-funded efficiency bet to a line item with cash flow behind it, sharply lowering the financial barrier to adoption. It also runs straight into the concern researchers raised in STAT's same-day piece: a per-use top-up rewards volume, and when it expires in two to three years the product must justify itself on clinical benefit alone.
Aidoc (Tel Aviv, Israel) is among the market leaders in imaging triage AI and in January 2026 expanded its CT triage platform to 14 indications, 11 of them new. CARE is its foundation-model line. CMS has not published the per-case add-on amount and none is inferred here (original release).
Nurses want a seat at the clinical-AI table — from strikes to co-design
STAT reported on Aug 11 (by Katie Palmer) that US nursing's pushback on clinical AI is becoming institutional. Two flashpoints are named: laid-off nurses at Montefiore in the Bronx questioning administrative AI displacing staff, and nurses at Kaiser Permanente in California striking over AI surveillance of their work. National Nurses United, representing more than 200,000 nurses, is the main actor.
Clinical AI increasingly fails not because the model is inaccurate but because nobody uses it as designed after deployment. Nurses are the frontline users of bedside alerts, deterioration detection and documentation AI — and among the least studied. Story 7 below makes the point concretely: in the UVA deterioration trial, clinicians moved 11% of patients between display and non-display beds, letting clinical judgement override the randomisation outright. STAT's argument is to convert an adversarial relationship into co-design — which is why "who was involved in deployment" is becoming as consequential a variable as model performance.
Five Nordic countries plus Estonia build a shared AI-Health platform on population registries
Ole A. Andreassen of the University of Oslo and 21 co-authors set out the technical foundations, data assets and deployment roadmap of the Nordic AI-Health Initiative in a Nature Medicine Comment (Aug 14, 2026). Participating institutions span Norway, Denmark, Iceland, Sweden, Finland and Estonia, including Karolinska Institutet, the University of Helsinki and Amgen's deCODE genetics. The stated aim is to "enable secure, regulation-compliant access and to deliver generalizable models for responsible AI-driven discovery in medicine."
Nordic longitudinal population registries — cradle-to-grave, multi-generational, linkable — are a globally scarce asset that GDPR and cross-border governance have made hard to use jointly. The significance is not any single model but that this is among the first regional attempts, post-EU AI Act, to build compliance into the data infrastructure itself rather than bolt it on afterwards. The reference value for Taiwan is high: the NHI database has the same longitudinal advantage and the same cross-institutional governance bottleneck (see the Taiwan Angle below). Note this is a Comment, not primary research — no cohort sizes or dated milestones are published.
RSNA opens a $77,000 knee-MRI AI challenge — the first to pair images with reports in 12 languages
RSNA announced the 2026 Knee Abnormality Detection AI Challenge on Aug 5, 2026. The dataset holds more than 5,000 knee MRI exams from 16 sites worldwide, with radiology reports in 12 languages. Total prize money is $77,000, including awards for the most efficient models; the challenge runs on Kaggle through Oct 22, winners announced in November and recognised at RSNA 2026 (Nov 29–Dec 3, Chicago). Co-leads are Po-Hao Chen, MD, MBA and Naveen Subhas, MD, MPH of Cleveland Clinic (AuntMinnie).
RSNA calls it "the most real-world challenge yet" precisely because it withholds clean structured labels and forces entrants to learn from messy, variably formatted, multilingual reports. That matters disproportionately for non-anglophone health systems, Taiwan included: most imaging-AI training corpora and benchmarks are English-report-based, and language mismatch is the most commonly underestimated source of performance drop at deployment. The second signal is the efficiency prize — academic benchmarks are finally treating inference cost as a first-class metric rather than optimising AUC alone.
A 10,422-visit deterioration-detection cluster RCT: no difference in the primary outcome — and the randomisation broke
Jessica Keim-Malpass, Jamieson M. Bourque and colleagues at the University of Virginia School of Medicine published a cluster-randomised controlled trial in Scientific Reports (Feb 5, 2026) covering 10,422 inpatient visits on an 85-bed cardiology and cardiac surgery ward. The intervention was CoMET (Continuous Monitoring of Event Trajectories) from Nihon Kohden Digital Health Solutions, which passively displays a risk trajectory refreshed every 15 minutes from machine-learning models over cardiorespiratory monitoring and vital signs. The primary outcome was hours free of clinical deterioration (death, emergent ICU transfer, emergent intubation, cardiac arrest, emergent surgery) within 21 days of admission.
"There was no change in the primary outcome between groups" — passively displaying a risk score did not change patient outcomes. Two methodological details deserve more attention than the headline. Only 5.3% of patients had a deterioration event, so power was structurally thin; and clinicians moved 11% of patients between display and non-display beds, preferentially shifting sicker patients into monitored beds, compromising the randomised design outright. This is exactly the bind flagged in Nature's July 28 comment, "Medical AI has a measurement problem" (Arjun K. Manrai, Harvard Medical School): evidence-based medicine assumes the yardstick is reliable, and in the AI setting that assumption is often shaky. Clinical AI that only displays and does not redesign the workflow and escalation path tends to produce precisely this result.
02 — Product Analysis
Aidoc CARE Body CT Multi-Triage
Imaging triage · healthcare foundation model · Aidoc (Israel)
Function and position. Detects multiple suspected urgent conditions on chest/abdomen/pelvis CT, with and without contrast, and reorders the worklist — "look at the right scan first," not autonomous interpretation (AuntMinnie).
- Strength : both the regulatory and payment tracks are complete — Breakthrough designation Sept 2025 → FDA clearance Jan 2026 → NTAP effective Oct 1, 2026, and the three-year add-on window gives hospitals a predictable financial model.
- Strength : one foundation model spanning many indications dilutes the "one algorithm, one submission" cost structure; the platform now carries 14 CT triage indications.
- Concern : triage endpoints are usually turnaround time, not mortality or readmission; researchers quoted by STAT note a per-use top-up structurally rewards volume.
- Concern : CMS has not published the per-case amount or an expected spend ceiling (Radiology Business), so hospital ROI cannot be computed from public data.
Retina4IRD
Rare eye disease · multimodal decision support · Shanghai Jiao Tong University (academic)
Function and position. Integrates multimodal retinal imaging to rank candidate causative genes for suspected inherited retinal disease, guiding the clinician's subsequent genetic testing and management (Nature Medicine).
- Strength : one of very few medical-AI trials that measures management decisions, not just accuracy — composite downstream score 37.7 vs 28.5, p < 0.001.
- Strength : validated across China, South Korea and Poland, with 85.6% external top-5 accuracy against 90.4% internally — better generalisation evidence than most single-country studies.
- Concern : top-1 accuracy is only 37.8% (vs 22.4%) — in practice it shortens the shortlist rather than delivering an answer, and expectations should be set accordingly.
- Concern : the randomised cohort is only 300 people (295 analysed), with no public commercialisation or regulatory pathway — deployment remains distant.
03 — Companies & Competition
| Company / system | Evidence | Position |
|---|---|---|
| Aidoc CARE Body CT Multi-Triage |
Complete regulatory and payment path (breakthrough Sept 2025, clearance Jan 2026, NTAP Oct 1, 2026); public prospective outcome data is limited. | Platform play across indications — 14 CT triage indications; the moat is deployment scale plus first-mover status on reimbursement. |
| ScreenPoint Medical Transpara v1.7 |
The strongest set: a 31,301-woman prospective paired non-inferiority trial (Nature Medicine, Mar 19, 2026) — 63.6% fewer radiologist readings, cancer detection +15.2% (6.3→7.3 per 1,000), recall +14.8% (4.8%→5.5%). | Deep in a single indication (breast screening), winning on randomised evidence, with the Swedish MASAI trial's interval-cancer and sensitivity data behind it. |
| Nihon Kohden CoMET |
Has completed a large cluster RCT, but the result was negative — 10,422 visits, no primary-outcome difference, with 11% of patients moved across arms. | A bedside-monitoring hardware vendor extending into predictive analytics: distribution through installed devices is the advantage, passive display is the design weakness. |
| Scanslated patient-friendly reports |
Large real-world usage data: 391,713 exams at Stanford and Duke children's hospitals (Jan 2024–Nov 2025) — 92.6% of families reported better comprehension, but only 9% opened the patient-friendly version versus 50.5% for the standard report. | Attacks the neglected patient-comprehension layer; the hospital pitch is satisfaction and retention (91.8% of families more likely to return) rather than diagnostic performance. |
04 — Taiwan Angle
(1) The multilingual-report challenge is an opening for Taiwanese datasets. RSNA's knee-MRI challenge is the first to include radiology reports in 12 languages from 16 institutions, an implicit admission that non-English reporting is the real-world norm. Taiwan's mixed Chinese-English reports — English structured fields, Chinese narrative — are essentially absent from international benchmarks, and language mismatch is a common source of the performance drop seen when imported imaging AI is deployed locally. Entering the challenge, or building a local multilingual benchmark, is an actionable move either way.
(2) The Nordic platform's governance design is more worth copying than its models. Taiwan's NHI database shares the Nordic registries' longitudinal scarcity value — and their bottleneck: cross-institutional governance and compliance. The Nordic AI-Health Initiative (Aug 14) writes compliance into the infrastructure rather than bolting it on, which is a ready-made comparator for MOHW's Guidelines on Generative AI Use in Healthcare Institutions and the existing three national smart-healthcare centres, including the Clinical AI Validation Center.
(3) Payment is the real adoption gate — and Taiwan's box is still empty. Aidoc's NTAP means US hospitals now have cash flow behind imaging-AI adoption. Taiwan's NHI still has no corresponding AI-software payment category, so deployment costs sit almost entirely with hospitals. Health Minister Shih Chung-liang set out a "333 policy" and an NT$48.9 billion budget at the Kaohsiung Medical University forum (June 27, 2026), targeting record interoperability across academic medical centres by year-end and regional and district hospitals within two years; the same event highlighted KMU's NVIDIA-based colorectal cancer detection system built with Foxconn. Interoperability is necessary — but without a payment design behind it, the Aidoc pattern of "adequate evidence, viable finances" stays hard to replicate here.
05 — Further Reading
-
Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people — Nature Medicine, 2026/8/4
The one to read end-to-end. Open access, with the full 4×2 factorial design and every subgroup analysis; anyone designing a clinical AI interface should start with the methods and Figure 3.
-
Medical AI has a measurement problem — Nature, 2026/7/28
Arjun K. Manrai's short comment states the core problem most clearly — evidence-based medicine assumes a reliable yardstick, and AI undermines that assumption. It is the frame for this year's run of null trials.
-
AI-based triage and decision support in mammography and digital tomosynthesis: a paired, noninferiority trial — Nature Medicine, 2026/3/19
A 31,301-woman prospective trial and the hardest evidence yet that imaging AI cuts reading volume — but recall also rose 14.8%, which belongs in the same sentence as the benefit.
-
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks — Nature Medicine
If general models keep winning on benchmarks, what exactly is the moat for narrow clinical AI? Required background for any medical-AI product strategy discussion.
-
Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial — Nature Medicine
Pragmatic cluster randomisation is the closest thing to a real-world test of clinical AI. Read alongside the UVA CoMET trial to see whether the bottleneck is intervention design or model performance.
06 — References
- Xu, X., Hu, H., … Ghassemi, M. “Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people.” Nature Medicine, vol. 32 no. 8, 2026-08-04. nature.com/articles/s41591-026-04553-w
- Jia, H., Qian, B., Sun, X. “AI-based clinician decision support system for diagnosis of inherited retinal diseases: a multicenter, randomized trial.” Nature Medicine, 2026-07-24. nature.com/articles/s41591-026-04545-w
- “CMS approves Medicare add-on payment for Aidoc CT triage AI.” AuntMinnie, 2026-08-13. auntminnie.com
- “Aidoc's CARE Body CT Multi-Triage receives eligibility for Medicare New Technology Add-on Payment.” PR Newswire (Aidoc press release). prnewswire.com
- “Medicare approves new technology add-on payment for inpatient radiology AI solution.” Radiology Business. radiologybusiness.com · “CT Triage Software Garners NTAP Reimbursement from CMS.” Diagnostic Imaging. diagnosticimaging.com
- Palmer, K. “Nurses seek a seat at the table as they fight expanding clinical AI.” STAT, 2026-08-11 (STAT+ paywall). statnews.com
- Palmer, K. “What Medicare incentives for AI-based devices mean for tech companies — and hospitals.” STAT, 2026-08-13 (STAT+ paywall). statnews.com
- Andreassen, O. A., Heimer, H., Kallioniemi, O., et al. “An AI-Health infrastructure for the Nordic region: technical foundations, data assets, and a roadmap for deployment.” Nature Medicine (Comment), 2026-08-14. nature.com/articles/s41591-026-04575-4
- “RSNA Launches Knee Abnormality Detection AI Challenge.” RSNA News, 2026-08-05. rsna.org · challenge page: rsna.org/artificial-intelligence · AuntMinnie coverage: auntminnie.com
- Keim-Malpass, J., Bourque, J. M., et al. “A randomized controlled trial of artificial intelligence-based analytics for clinical deterioration.” Scientific Reports, 2026-02-05. nature.com/articles/s41598-026-39051-z
- Manrai, A. K. “Medical AI has a measurement problem.” Nature, 2026-07-28. nature.com/articles/d41586-026-02125-z
- Álvarez-Benito, M., et al. “AI-based triage and decision support in mammography and digital tomosynthesis for breast cancer screening: a paired, noninferiority trial.” Nature Medicine, 2026-03-19. nature.com/articles/s41591-026-04277-x
- “Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading (MASAI).” The Lancet. thelancet.com
- “General-purpose large language models outperform specialized clinical AI tools on medical benchmarks.” Nature Medicine. nature.com/articles/s41591-026-04431-5 · “Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial.” Nature Medicine. nature.com/articles/s41591-026-04503-6
- Johnston, A., et al. “Patient-friendly radiology reports improve family comprehension” (Pediatric Radiology, reported 2026-08-03). AuntMinnie. auntminnie.com
- 「高醫大論壇揭示AI醫療新局!衛福部推『333政策』」,聯合新聞網,2026-06-27。udn.com ·「衛生福利部頒布『醫療機構應用生成式人工智慧指引』」,理律法律事務所。leeandli.com · 臺灣智慧醫療三大中心 aicenter.mohw.gov.tw
- “FDA clears CT-based AI triage platform Aidoc.” Diagnostic Imaging, 2026-01-21. diagnosticimaging.com