◆ AI & Medical AI Daily
–
Monday · Clinical Applications & Research

Same day, same journal, two endings: AI reading CT recovered 51 missed lesions; AI reading charts was quietly abandoned

On August 19, Nature Medicine published two prospective clinical evaluations on the same day, and they point in opposite directions. One is LiON, a contrast-CT liver AI from Shengjing Hospital of China Medical University with Alibaba's DAMO Academy and others: AUC 0.952 in a 10,333-patient prospective trial, and in live workflow it recovered 51 previously overlooked lesions, 15 of them malignant, prompting 37 amended reports. The other is SHAKED, an LLM decision-support system in the emergency department at Rambam Health Care Campus in Israel: 1,138 patients, zero adverse events, 99 of 100 sampled outputs judged clinically appropriate by expert review — and yet clinician uptake slid from 68% to 30%, length of stay was 4.9 hours in both wings, and the authors concluded plainly that the findings "do not justify clinical deployment at this stage." Read side by side, and against the 1,357-to-3 denominator PLOS Digital Health published last week, today's question is singular: is clinical AI decided by accuracy, or by whether anyone keeps using it?

01 — Top Stories

Seven stories: from two Nature Medicine papers to an endoscopy AI switched on in Singapore
Story of the day Nature MedicineRambamDECIDE-AI8/19

SHAKED, an emergency-department LLM: safe, accurate, and gradually abandoned

What

On August 19, Rambam Health Care Campus in Israel — with collaborators at the Technion, Mayo Clinic and Beth Israel Deaconess — published a DECIDE-AI stage-1 prospective pilot in Nature Medicine. The LLM clinical decision support system SHAKED ran in one emergency-department wing while a parallel wing kept standard operations, over four weeks, with 1,138 patients analysed (584 in the intervention wing). The result splits three ways. (1) It was safe: no adverse events were detected and expert review rated 99 of 100 sampled outputs clinically appropriate. (2) The benefit did not materialise: length of stay was 4.9 hours in both wings (P = 0.99) and the 9.4-minute reduction in consultation cycle time missed significance (P = 0.077). (3) Uptake collapsed: adoption fell from 68% to 30% across the study and tracked workload closely (OR = 0.72 per shift hour, 95% CI 0.62–0.83). The one setting that bucked the trend was radiology consultation (OR = 2.98, 95% CI 1.58–5.63). The authors' conclusion is unusually blunt: "sustained clinician engagement, rather than algorithmic accuracy, may be the key barrier," and the findings "do not justify clinical deployment of AI clinical decision support at this stage."

Why it matters

For three years nearly every clinical-LLM study has answered one question: is the model accurate enough? SHAKED replaces the question. The model was accurate (99/100) and safe (zero adverse events), yet in a live emergency department it was used less each week — and least by the clinicians who were busiest. That inverts the commercial story about AI saving time: if AI only gets used when you are not busy, it is not addressing busyness at all. OR = 0.72 per shift hour is the coefficient of the week, because it turns adoption into something measurable and usable as an endpoint. It also annotates the PLOS Digital Health finding that only 3 of 1,357 cleared devices were tested on patient outcomes: even when you do run the trial, you may measure nothing — because the intervention was diluted by its own users mid-study. Catching exactly this is what the DECIDE-AI framework was designed for, and SHAKED is one of the few studies to follow it honestly and publish the negative result.

Players

SHAKED is an in-house LLM decision-support system at Rambam (Haifa); authors Leibovitch, Ahituv, Gorenshtein, Aran, Sorka, Miron and Shelly are affiliated with Rambam, the Technion, Mayo Clinic and Beth Israel Deaconess. The underlying foundation model is not disclosed. The study is DECIDE-AI stage 1 (early live clinical evaluation), not an efficacy trial; the comparator is standard operations in a parallel wing of the same department.

Imaging Nature MedicineDAMO Academy盛京醫院8/19

LiON: a liver-cancer CT AI validated in 33,000 patients, catching 51 lesions the workflow had missed

What

The second Nature Medicine paper of the same day comes from Yu Shi (corresponding author) in the radiology department of Shengjing Hospital, China Medical University, with Alibaba's DAMO Academy, King's College London and EURECOM among institutions across China, France and the UK: a multicentre study plus single-arm clinical trial of the Liver DiagnOsis Network (LiON). Scale is the story: 6,443 patients for training, 22,251 for multicentre retrospective validation, 10,333 in the prospective trial. Performance: AUC 0.975 (95% CI 0.971–0.979) for malignancy in retrospective validation, 0.971 in the hepatic-steatosis subgroup and 0.924 in cirrhosis; the prospective primary endpoint was AUC 0.952 (95% CI 0.942–0.961). The weightier result is secondary: inside routine radiology workflow LiON surfaced 51 previously overlooked lesions including 15 malignancies, triggered 37 amended reports and led to 22 multidisciplinary-team escalations. It reads multiphase contrast-enhanced CT with flexible phase input (non-contrast, arterial, venous, optional delayed), outputs a patient-level malignancy score, segments eight lesion types, and folds in demographics, biomarkers and history.

Why it matters

LiON against SHAKED is close to a natural experiment. Both ran inside live clinical workflow, neither is a randomised efficacy trial, and only one produced identifiable patient-level consequences: 51 lesions, 37 amended reports, 22 MDT escalations. The difference is not model intelligence but the shape of the intervention. LiON exists as a second reader and safety net; its work happens outside the clinician's attention and requires no one to open anything. SHAKED asks a physician to take one more action at their busiest moment. That yields today's most usable design principle: adoption of clinical AI is inversely proportional to the active attention it demands from a clinician. The limits are equally clear: a single-arm trial has no comparator, so it cannot say whether those 51 lesions would eventually have been found without AI, nor convert them into survival; the authors themselves call for prospective comparative studies across health systems. LiON still sits outside that PLOS 0.2% — just nearer the threshold than most.

Players

Led by the radiology department at Shengjing Hospital, China Medical University (Shenyang), with Alibaba DAMO Academy as technical partner — the same group behind the DAMO PANDA pancreatic-cancer screening model, making liver its second organ-scale target. Academic partners: King's College London and EURECOM. Competitive frame: Aidoc (abdominal CT triage), United Imaging Intelligence, Infervision, and Lunit and DeepHealth concentrated in chest and breast — multiphase liver CT is comparatively uncrowded.

Evidence PLOS Digital HealthUniversity of TorontoMIT8/19–20

The denominator behind today's two papers: 1,357 cleared, 3 ever tested on patient outcomes

What

Rawan Abulibdeh (University of Toronto) with co-authors including MIT's Laboratory for Computational Physiology published on August 19 in PLOS Digital Health an audit of all 1,357 FDA-authorised AI medical devices through December 5, 2025: just 34 appeared in a registered clinical trial, 12 posted results, 12 reached peer review, and only 3 were tested against patient-centred outcomes such as mortality, stroke, hospitalisation or quality of life. As News-Medical summarised on Aug 20, the research was conducted almost entirely in well-resourced health systems and systematically excluded pregnant women, adults over 75 and non-English speakers; the authors warn the framework could amplify existing disparities and make low- and middle-income countries unwitting test populations for under-studied devices. Healio (Aug 21) and AuntMinnie also covered it.

Why it matters

We covered this study in yesterday's weekly review; it returns today because it is the denominator for the first two stories. Place LiON and SHAKED inside the framework and something easy to miss appears: neither paper would count among the 3 — LiON is single-arm, SHAKED measured adoption and process metrics, and neither used a patient outcome as its primary endpoint. Even the two most rigorous prospective studies of the week remain some distance from proving patients live longer or better. That is not a failure of the investigators but of the evaluation infrastructure for clinical AI itself: no standard endpoints, no established convention for cross-institution comparators, no tradition of treating adoption as a formal endpoint. The structure is what to remember today, not any single number.

Players

The unit of analysis is the whole FDA AI/ML-Enabled Medical Device List; no vendor is named. Taiwan's counterpart is the TFDA AI/ML medical device information and matchmaking platform, for which no equivalent validation audit exists.

Background study NEJM AIRCTNCT06963957

Why "the physician will catch it" is not a safety design: AI-trained doctors still lost 14 points to bad advice

What

To see why "99 of 100 appropriate" in story 1 is not itself a safety guarantee, it helps to read an earlier randomised trial. "Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy" in NEJM AI (May 2026 issue; NCT06963957) randomised 44 physicians who had each completed 20 hours of AI-literacy training to receive either error-free or deliberately flawed LLM recommendations. Diagnostic reasoning scores fell from 84.9% to 73.3% (adjusted difference −14.0 percentage points), and top-choice diagnostic accuracy from 90.5% to 76.1% (adjusted −18.3 points). The preprint is on medRxiv, and a follow-up RCT preprint from June 2026 tests behavioural nudges as mitigation.

Why it matters

Nearly every risk narrative for clinical AI rests on one assumption: a physician makes the final call, so AI errors get caught. This trial shows the assumption fails even among specifically trained physicians — and the degradation (−14.0 and −18.3 points) is large enough to cancel the gains most AI tools claim. Connect it back to story 1: SHAKED's "99 of 100 appropriate" means 1 in 100 is not appropriate, and this RCT says that one will not be reliably intercepted. It is also why LiON's safety-net posture in story 2 is structurally safer: AI that says "you may have missed something" costs a second look when wrong; AI that says "here is what you should do" costs a persuaded clinician. Caveat: 44 physicians in a single country (Pakistan) in a simulated-case design — a mid-sized study whose generalisation should stay conservative.

Players

No commercial product is involved; a general-purpose LLM supplied the recommendations, with errors deliberately injected to test independent clinical judgement. This kind of adversarial-advice stress test is not yet a pre-market requirement at any regulator — one of the gaps in the generative-AI device pathway the FDA is currently weighing.

Deployment AinexNUH Singapore8/18

Endoscopy AI arrives at Singapore's NUH: Korea's ENAD claims a 16% lift in adenoma detection

What

On August 18, Healthcare IT News reported that South Korean company Ainex has deployed its real-time endoscopy AI ENAD (Endoscopy as AI-powered Device) at National University Hospital (NUH) in Singapore for upper and lower gastrointestinal endoscopy, detecting lesions in real time and automatically characterising them to support clinical decisions. Ainex says peer-reviewed studies show a 16% increase in adenoma detection rate (ADR), and that ENAD runs in more than 200 health settings in South Korea with approvals in Thailand, Malaysia and Colombia.

Why it matters

Endoscopy AI is the closest thing clinical AI has to a proven category: ADR is an accepted colonoscopy quality metric with a documented link to subsequent colorectal cancer incidence, and computer-aided detection (CADe) has accumulated multiple randomised trials. It is also a textbook case of story 2's design principle — the AI draws a box on the screen and no clinician has to summon it, so nothing resembling SHAKED's decay in usage occurs. Two honest annotations: (1) the 16% is a vendor-aggregated figure and the report does not link the underlying paper; (2) recent literature notes CADe raises ADR partly by removing more tiny, low-risk adenomas, leaving long-term benefit contested. The market signal is more direct: clinical-AI procurement in Korea and Singapore has entered its cross-border export phase, and Taiwanese vendors face exactly this cohort of competitors.

Players

Ainex (South Korea); customer: National University Hospital, Singapore. Category competitors: Medtronic GI Genius, Fujifilm CAD EYE, Olympus ENDO-AID, and Korean peers. Taiwanese endoscopy-AI groups remain largely at the hospital-research-collaboration stage, with no cross-border rollout at comparable scale.

Debate JAMAAMAKhosla8/21

"AI ready for real-world cognitive tasks by 2030": Emanuel and Khosla against the AMA

What

On August 21 Fierce Healthcare reported on an escalating argument: bioethicist and oncologist Ezekiel J. Emanuel, with Neal Khosla of Curai Health and Vinod Khosla of Khosla Ventures, co-authored a JAMA viewpoint arguing that after reviewing all published work on AI in medicine since January 1, 2024, AI now matches or outperforms physicians on five core cognitive tasks — gathering patient information, diagnosing, selecting tests, recommending treatment and managing chronic disease — and will be ready for autonomous deployment in some, perhaps many, workflows by around 2030. They go further: physician oversight introduces errors and degrades otherwise superior AI performance. AMA CEO John Whyte counters that medicine requires human judgement beyond task execution, likening it to pilots remaining essential in the age of autopilot, and that empathy, communication and shared decision-making cannot be delegated to an algorithm.

Why it matters

The debate belongs beside stories 1 and 4 because both camps can cite today's evidence, and each cites only half of it. Emanuel and Khosla can point at story 4 and say physician oversight does introduce error — the RCT does show clinicians pulled off course by bad advice. But that RCT's premise is that the AI gave the wrong answer first, which is precisely the case for not removing oversight. The AMA can point at story 1 and say clinicians do not buy it — but SHAKED's decaying uptake does not prove the AI was worthless, only that it was designed to require effort. What both sides skip is story 2's answer: arguing over whether AI should replace physicians may be the wrong level of the problem. LiON neither replaced the radiologist nor waited to be opened by one; it simply changed the probability that a lesion is seen.

Players

Curai Health (AI-first virtual care, backed by Khosla Ventures); the American Medical Association; the University of Pennsylvania (Emanuel). Note the commercial backdrop: of the authors arguing for autonomous AI, one runs an autonomous-care company and one is its investor, so the JAMA viewpoint should be read as an interested position paper rather than a neutral literature review.

Care model HIMSS26 APACSingapore8/22

In dementia care, detection is not the bottleneck: the 39% MCI window is stuck between three disconnected systems

Why it matters

The first six stories ask whether AI is accurate and whether anyone uses it; this one moves up a level: even a perfect screening AI changes nothing if there is no service to hand the positive case to, and no record coming back. That maps onto Taiwan closely — Taiwan is not behind on dementia screening instruments, digitised cognitive scales or imaging biomarkers, yet the path from community centres through primary care to medical centres that turns "we found it" into "they are being treated, and it is recorded" is equally missing. In story 1's vocabulary this is adoption collapse in another form: not clinicians abandoning the tool, but detected patients falling through the gaps between systems. Caveats: this is an interview with an interested vendor; Brain Health Playground has no peer-reviewed outcome evidence cited, and the 39%/22% figures come from the referenced meta-analysis, not from the company's product.

Players

Youngand Inc. (South Korea; Brain Health Playground), operating in Singapore and across APAC. Read alongside Healthcare IT News on Aug 21 on how rural hospitals prioritise AI investment — both turn on which tier of institution the AI capability lands in.

02 — Product Analysis

Two clinical AI systems published the same day: one stands behind the clinician, one stands in front

LiON(Liver DiagnOsis Network)

Contrast-CT liver malignancy AI · academic + DAMO Academy · China

Function and position. Takes multiphase contrast-enhanced CT (non-contrast, arterial, venous, optional delayed), outputs a patient-level malignancy score, segments eight liver lesion types and integrates demographics, biomarkers and history. It is positioned not as a replacement for the radiologist but as an additional reader and diagnostic safety net (Nature Medicine, 2026-08-19).

  • Strength : scale and subgroup robustness. AUC 0.975 across 22,251 retrospective patients and 0.952 in a 10,333-patient prospective trial, holding up in the hard subgroups — 0.971 in hepatic steatosis and 0.924 in cirrhosis. Subgroup collapse is where imaging AI usually fails.
  • Strength : identifiable clinical consequences, not just statistics. 51 overlooked lesions (15 malignant), 37 amended reports, 22 MDT escalations — something most AUC papers cannot produce, and the reason it sits closer to patient outcomes than story 1 does.
  • Concern : a single-arm design cannot answer the counterfactual. Without a comparator there is no way to know whether those 51 lesions would have surfaced later anyway, or how much later, and no path to survival or stage-shift. The authors themselves call for prospective comparative studies across health systems.
  • Concern : geographic concentration of training and validation. The data is predominantly from Chinese health systems, where hepatitis-B-driven hepatocellular carcinoma prevalence differs from Western populations. Taiwan shares the hepatitis-B-dominant aetiology, so transfer is more plausible than to the US or EU — but local validation is still required. This is story 3's "68% of trials run in one country" problem, with a different country.

SHAKED

Emergency-department LLM decision support · built in-house · Israel

Function and position. LLM-based clinical decision support inside emergency-department workflow, including consultation scenarios. It is an assistive layer the clinician actively queries — and that word "actively" is its fundamental difference from LiON (Nature Medicine, 2026-08-19).

  • Strength : the study's honesty is its greatest asset. DECIDE-AI stage 1, a parallel control wing, negative results published, and an explicit statement that the findings "do not justify clinical deployment at this stage." In a field saturated with vendor-reported numbers, a paper willing to disconfirm itself is worth more than most positive ones.
  • Strength : it located a sub-scenario with genuine demand. Overall uptake fell, but radiology consultations ran at OR = 2.98 (95% CI 1.58–5.63). That is a viable product direction: stop building the assistant that answers anything, and build the one scenario a clinician will go out of their way to use.
  • Concern : uptake correlating negatively with workload is a design-level flaw. Odds of use fall to 0.72 per additional shift hour. A tool that claims to reduce burden goes unused precisely when burden peaks — meaning the intervention point is wrong and should move to where no clinician action is required (pre-generated, passively surfaced).
  • Concern : the residual risk in "99 of 100" is unaddressed. Read against story 4's automation-bias RCT (flawed advice cut diagnostic reasoning by 14.0 points), that 1% will not be reliably intercepted. Four weeks, one hospital, and an undisclosed base model also make the result hard to transfer to other systems or model versions.

03 — Companies & Competition

Six players from today, and where each places AI inside the clinical workflow
Organisation Latest & numbers Workflow position & moat
阿里巴巴達摩院
DAMO Academy · imaging models
Co-published LiON with Shengjing Hospital: 6,443 training / 22,251 retrospective / 10,333 prospective, AUC 0.975 and 0.952, 51 overlooked lesions recovered; previously published a pancreatic-cancer screening model. Position: passive safety net (second reader). Moat = large multicentre Chinese imaging data × accumulated organ-scale models. Weakness: geographic concentration; no FDA/CE pathway disclosed.
Rambam Health Care Campus
Haifa, Israel · in-house LLM CDS
SHAKED ED pilot: 1,138 patients, zero adverse events, 99/100 outputs appropriate, uptake 68%→30%, LOS 4.9h with no difference; collaborators include the Technion, Mayo Clinic and Beth Israel Deaconess. Position: active query layer. Moat = in-house data and trial-execution capability, not a product. Its export is methodological: turning adoption into a measurable endpoint.
Ainex
South Korea · real-time endoscopy AI
ENAD deployed at NUH Singapore; claims a 16% ADR lift; 200+ Korean sites plus approvals in Thailand, Malaysia and Colombia (Aug 18). Position: real-time passive prompting (on-screen boxes). Moat = endoscopy-vendor channels × stacked multi-country approvals. Competitors: Medtronic GI Genius, Fujifilm CAD EYE, Olympus ENDO-AID.
長佳智能 EverFortune.AI
Taiwan (6841) · imaging AI devices
Appendicitis CT AI cleared by Taiwan's TFDA (announced Aug 18, 2026), following a US FDA 510(k) in April; 57 device authorisations in total (15 US FDA, 24 TFDA) and 70+ customers. Position: acute triage prompting. Moat = stacked Taiwan-plus-US regulatory credentials × hospital channels. Weakness: like LiON and Aidoc, no patient-outcome-level evidence, so competition reverts to distribution and price.
DeepHealth(RadNet)
US · mammography AI workflow
A July 23 Radiology study: 95 radiologists, 109 facilities, ~578,000 screening mammograms; generalists' cancer detection rose from 3.76 to 4.99 per 1,000, approaching specialists' 4.76; PPV of recalls up 15.09%, recall rate also up 14.79%. Position: workflow redesign, not just an algorithm. Moat = owning the imaging-centre network, hence control of both workflow and data. Its pitch — closing the subspecialist gap — speaks directly to today's theme.
Curai Health / Khosla Ventures
US · autonomous-AI advocacy
With Emanuel in JAMA, argues AI already matches or beats physicians on five cognitive tasks, is deployable autonomously by 2030, and that physician oversight introduces error; AMA CEO John Whyte publicly disagrees (Aug 21). Position: replacement layer. Moat = capital and narrative rather than clinical evidence. Note the authors are simultaneously operator and investor in the category: an interested position paper.

04 — Taiwan Angle

Taiwan collected its 57th device authorisation this week: the next question is not how many, but whether anyone uses them

Taiwan's item of the day. On August 18, EverFortune.AI (TWSE: 6841) announced that its appendicitis AI software, the EFAI ERSUITE CT Appendicitis Assessment System, had received Taiwan TFDA clearance, following its US FDA 510(k) in April and completing coverage of both markets; the company now holds 57 device authorisations worldwide (15 US FDA, 24 Taiwan TFDA) across 70-plus customers. The product applies deep learning to contrast-enhanced abdominal CT and flags suspected adult appendicitis to the clinical team. Read alongside it: on August 21 TÜV Rheinland completed third-party testing of EverBot Technology's patient-education assistant "EirBot," reporting a 100% retrieval hit rate across disease, surgery, examination, medication and care-process queries, 99.9% accuracy and faithfulness in its patient-education answers, and 100% classification accuracy on security threats and high-risk questions, with methods referencing ISO/IEC TS 4213:2022.

Converting today's international news into Taiwan's problem. Taiwan is not slow on AI devices: the TFDA runs an AI/ML device information and matchmaking platform, EverFortune alone holds 24 TFDA authorisations, and the MOHW has stood up three national smart-healthcare AI centres. But SHAKED in story 1 raises a question Taiwan has not systematically answered: of the AI already cleared and already installed, what share is still in use three months later? There is no public tracking of clinical-AI utilisation in Taiwan — not in hospital self-assessment, not on the NHI claims side, not in TFDA post-market surveillance, none of which treat adoption as a required reporting item. And SHAKED's central finding is exactly that adoption declines with workload, and faster than anyone expects. Taiwan's clinical staffing is already stretched thin, so that coefficient will be steeper here, not flatter.

Three actionable directions. (1) Make utilisation a post-market surveillance item. Story 3 shows only 3 of the FDA's 1,357 devices were tested on patient outcomes; closing that evidence gap quickly is unrealistic for Taiwan, but adding the utilisation layer is cheap and immediately feasible — it is a precondition for efficacy and among the hardest metrics to fake. (2) Prioritise passive AI in procurement. Today's clearest rule of thumb: AI that requires no clinician to open it — LiON, ENAD — sustains its usage, while AI that asks for one more click decays. For hospitals ranking a limited budget, that predicts more than AUC does. (3) Liver AI deserves local validation. Taiwan shares China's hepatitis-B-dominant hepatocellular carcinoma profile, so LiON's aetiological base is closer to Taiwan's than most Western imaging AI, making local prospective validation a better bet — and hepatology is already one of the few areas where Taiwan's clinical research base is internationally competitive.

The scepticism worth keeping. EirBot's 100%/99.9% come from a single commissioned test with an unpublished question bank, sample size and difficulty distribution, and cannot be equated with trial-grade evidence; EverFortune's 57 authorisations are a regulatory-access number, not a clinical-benefit number — the gap between those two is what story 3's entire paper is about. Taiwan performs well at obtaining clearances. What today's seven stories add up to is that clearance is the qualifying round; the real contest starts on day 90 after installation.

05 — Further Reading

Five worth reading end to end
  1. Leibovitch, L., et al. “Prospective evaluation of a large language model clinical decision support system in the emergency department” — Nature Medicine (2026-08-19)

    Today's essential read. Pay particular attention to how the paper models declining adoption over time — treating usage behaviour as an outcome variable with a workload interaction. That method is both the scarcest thing in clinical-AI research and the easiest to replicate elsewhere. Any hospital preparing to deploy an LLM tool should draft its evaluation plan from this design.

  2. Shi, Y., et al. “Large-scale AI-guided liver malignancy diagnosis: multicenter study and a single-arm trial” — Nature Medicine (2026-08-19)

    One of the largest prospective imaging-AI studies to date. Beyond the AUC, the instructive part is how it reports lesions the AI surfaced and what happened next — amended reports, MDT escalations — a model for translating a statistic into clinical action. For any Taiwanese hepatology group planning local validation, this is a ready-made protocol template.

  3. “Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial” — NEJM AI (2026-05, NCT06963957)

    Not this week's news, but necessary background for every story above. Read it to dismantle an assumption too often treated as a premise: that a physician in the loop makes the system safe. If the full text is out of reach, the medRxiv preprint is free and carries the same figures and conclusions.

  4. Basu, S. & Huynh, B. “Mitigating hallucinations in healthcare AI: a systematic review of evidence-based strategies” — BMC Health Services Research 26:1115 (2026-06-11)

    Open access, 44 studies, and it tabulates effect sizes across seven mitigation strategies in comparable form: RAG cuts hallucinations 30–50%, knowledge graphs 25–45%, human-in-the-loop up to 95% but does not scale. For a hospital selecting an LLM architecture, this is the most practical decision aid currently available.

  5. Kömen, J., et al. “Towards robust foundation models for digital pathology” (PathoROB) — Nature Communications 17:5218 (2026-06-11)

    Best read against today's story 2. The benchmark subjects 20 pathology foundation models to non-biological variation such as scanner and staining differences, and finds that when training data correlates with medical centre, tumour detection accuracy falls from above 92% to 53–87% — the model learned the institution, not the biology. Anyone procuring AI across multiple hospitals should ask this question first.

06 — References

References
  1. Leibovitch, L., Ahituv, A., Gorenshtein, A., Aran, D., Sorka, M., Miron, K., Shelly, S. “Prospective evaluation of a large language model clinical decision support system in the emergency department.” Nature Medicine, 2026-08-19. nature.com/articles/s41591-026-04601-5
  2. Shi, Y., et al. “Large-scale AI-guided liver malignancy diagnosis: multicenter study and a single-arm trial.” Nature Medicine, 2026-08-19. nature.com/articles/s41591-026-04589-y · journal listing nature.com/nm/articles
  3. Abulibdeh, R., et al. “1,357 AI medical devices cleared, 3 actually tested on patient outcomes.” PLOS Digital Health, 2026-08-19. DOI: 10.1371/journal.pdig.0001597. journals.plos.org
  4. “FDA-cleared medical AI rarely tested for real-world benefits.” News-Medical, 2026-08-20. news-medical.net · “Most AI tools cleared by FDA were not tested on clinical outcomes.” Healio, 2026-08-21. healio.com · “Most FDA-cleared AI medical devices not tested on patient outcomes.” AuntMinnie. auntminnie.com
  5. U.S. FDA. “Artificial Intelligence-Enabled Medical Devices” (official device list). fda.gov
  6. “Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial.” NEJM AI, 2026-05 (Vol. 3, No. 5), NCT06963957. ai.nejm.org · summary coverage: bjh.be · preprint: medRxiv
  7. “Mitigating Automation Bias in Physician-LLM Diagnostic Reasoning Using Behavioral Nudges: A Randomized Controlled Trial.” medRxiv, 2026-06. medrxiv.org
  8. “Korean AI to support digestive diagnosis at NUH.” Healthcare IT News (Asia), 2026-08-18. healthcareitnews.com
  9. “As AI advances, a debate grows over the future role of doctors.” Fierce Healthcare, 2026-08-21. fiercehealthcare.com
  10. “In dementia care, detection is not the bottleneck.” Healthcare IT News (Asia), 2026-08-22. healthcareitnews.com · “Rural hospitals face high stakes in allocating AI investments.” Healthcare IT News, 2026-08-21. healthcareitnews.com
  11. “FDA weighs regulatory approach for genAI medical devices.” Healthcare IT News, 2026-08-21. healthcareitnews.com
  12. “AI-backed imaging workflow helps generalist radiologists perform like breast specialists.” Radiology Business, 2026-07-23(underlying study in RSNA's Radiology). radiologybusiness.com
  13. Basu, S., Huynh, B. “Mitigating hallucinations in healthcare AI: a systematic review of evidence-based strategies.” BMC Health Services Research 26:1115, 2026-06-11. link.springer.com · PubMed
  14. Kömen, J., Müller, K.-R., et al. “Towards robust foundation models for digital pathology” (PathoROB). Nature Communications 17:5218, 2026-06-11. nature.com/articles/s41467-026-73923-2
  15. 「長佳智能闌尾炎AI醫療器材軟體攻台美市場 累計有57張海內外醫材證」,台灣光鹽生物科技學苑,2026-08-19(company announcement Aug 18)。biotech-edu.com
  16. 「醫療 AI 回答不能『無中生有』 德國萊因完成長聯科技『愛寶』AI 衛教測試」,經濟日報,2026-08-21。money.udn.com · referenced standard ISO/IEC TS 4213:2022
  17. Taiwan device and policy background: TFDA 智慧醫療器材資訊暨媒合平台 · 衛福部臺灣智慧醫療三大中心
  18. Recent editions (for cross-reference): 8/20 技術 · 8/21 產品 · 8/22 廣義 AI · 8/23 本週回顧
Editor's note: (1) All figures in stories 1 and 2 are taken directly from the public abstracts and results sections of the two Nature Medicine papers; both full texts are paywalled and nothing beyond the verifiable abstract material is cited here. (2) The NEJM AI paper in story 4 was published in May 2026 and is not this week's news; it is included as necessary background for reading story 1's safety figures. Its full text requires a subscription, so figures here come from the public abstract and the free medRxiv preprint, which agree. (3) ENAD's "16% increase in adenoma detection rate" in story 5 is a vendor-aggregated figure; the source report links no underlying paper and this report could not verify it independently. (4) The JAMA viewpoint in story 6 is paywalled and is reported here via Fierce Healthcare's public coverage; two of its authors are respectively the operator and the investor of an autonomous-AI care company, making it an interested position paper. (5) Story 7 is an interview with an interested vendor; Brain Health Playground has no peer-reviewed outcome evidence, and the 39%/22% figures come from the meta-analysis it cites. (6) In the Taiwan section, EirBot's 100%/99.9% come from a single commissioned test whose question bank, sample size and difficulty distribution are unpublished and should not be equated with trial-grade evidence; EverFortune's authorisation count is company-reported via media and is a regulatory-access metric, not a clinical-benefit one. (7) Every link in this report is a URL actually visited during research; none was constructed or inferred.