◆ AI & Medical AI Daily
–
Monday · Clinical Applications & Research

This week's clinical evidence on medical AI all points one way — accuracy is no longer the bottleneck: Rambam's emergency-department LLM had 99 of 100 outputs judged clinically appropriate while physician uptake fell from 68% to 30% in four weeks; AI beat case managers at predicting discharge dates on the day of admission and lost by 2x in the final 24 hours; and a CD276 antibody-drug conjugate designed by 37,000 AI agents at Stanford was independently arrived at by Merck and granted FDA breakthrough designation

Line up the studies that entered clinical circulation this week and they turn out to be facets of a single finding. In Nature Medicine, a prospective trial of an LLM decision-support system at Israel's Rambam Health Care Campus split 1,138 emergency-department patients across two parallel wings; expert review rated 99 of 100 sampled outputs clinically appropriate, with no adverse events — and then physician uptake fell from 68% to 30% over four weeks, the odds of use dropping to 0.72 for each additional hour of shift workload. Length of stay was 4.9 hours in both wings, P = 0.99. The authors' own conclusion: the barrier is sustained clinician engagement, not algorithmic accuracy, and the evidence does not yet support deployment. JAMA Network Open supplies the other half, across 22,349 inpatient encounters: AI edged out case managers at predicting the discharge date on the day of admission (mean absolute error 4.20 days versus 4.27), then lost 1.93 to 0.98 days in the final 24 hours — it wins precisely when the prediction is least useful. NEJM AI, meanwhile, points out that the premise of all this is not yet secure: a patient-facing RAG chatbot gave up its system prompt, retrieval configuration, backend endpoints and 1,000 most recent patient conversations to nothing more than browser developer tools, with no authentication required at any point (arXiv:2605.00796). The one large story running the other way is the 37,000-agent "Virtual Biotech" from James Zou's lab at Stanford: it designed an antibody-drug conjugate against CD276 for lung cancer, and months later Merck independently arrived at the same therapeutic design and won FDA breakthrough designation for it (Stanford Medicine, 2026-09-17).

01 — Top Stories

Eight studies that entered clinical circulation this week, ordered by how close they get to the bedside
Prospective trial Rambam · Technion · MayoSHAKED8/19

The first prospective controlled evaluation of an emergency-department LLM: nearly every output passed, and physicians stopped using it within four weeks

What

Rambam Health Care Campus in Haifa, with collaborators at the Technion and Mayo Clinic, put an LLM-based clinical decision support system called SHAKED into one of two parallel-running emergency wings while the other kept standard practice, over four weeks, analysing 1,138 patients (584 allocated to the intervention wing) and reporting to DECIDE-AI stage 1 specification. Result: expert review rated 99 of 100 sampled outputs clinically appropriate, with no adverse events detected; but clinical adoption fell from 68% to 30%, driven mainly by workload-related disengagement (OR = 0.72 per shift hour), and the one use physicians clearly preferred was radiology consultation (OR = 2.98). Length of stay was 4.9 hours in both wings (P = 0.99), and the consultation cycle shortened by 9.4 minutes without reaching significance (P = 0.077). Nature Medicine, 2026-08-19 / PubMed 42618632

Why it matters

This is a rare case in medical AI where almost everything went right and the system still cannot be deployed — and the reason is not the model. For two years the standard argument has run from benchmark score to clinical value; this trial unpacks the middle of that inference. Even with output quality close to perfect, uptake collapses as soon as clinical workload rises, and once uptake collapses any downstream endpoint — length of stay, consultation efficiency — is diluted past the point of measurement. The authors put it plainly: the barrier is sustained clinician engagement, not algorithmic accuracy. For buyers the implication is that no demo accuracy figure substitutes for the number of people still using it in week three.

Discount this

Single centre, four weeks, and DECIDE-AI stage 1 is by definition an early-stage design, not a randomised efficacy trial. Length of stay and consultation cycle are secondary endpoints and the study is not powered to detect small effects, so "P = 0.99" does not mean "confirmed no effect". The 99-of-100 appropriateness rate comes from sampled expert review, not a full audit.

Drug discovery Stanford · James ZouVirtual Biotech · CD276 ADC9/17

A lung-cancer ADC designed by 37,000 AI agents was independently arrived at by Merck — and granted FDA breakthrough designation

What

The team of James Zou, associate professor of biomedical data science at Stanford, transplanted the org chart of a pharmaceutical company into a multi-agent system: separate divisions for target discovery, molecule design and clinical trials, orchestrated by an "AI chief scientific officer" agent, with the trials division alone running roughly 37,000 agents on top of an AI-native data platform called Paperclip (which the team says cuts time and cost by "over an order of magnitude"). One output was an antibody-drug conjugate (ADC) against CD276 for lung cancer; months later Merck independently developed and validated the same therapeutic design, which went on to receive FDA breakthrough designation. The team separately reports that targets selected on single-cell features were "about 50% more likely to reach market". Stanford Medicine, 2026-09-17 / VentureBeat / Nature news feature

Why it matters

The hard part of AI drug discovery is not generating hypotheses but showing that a hypothesis is not a hallucination. Everything that makes this story worth reporting rests on one clause: Merck converged independently on the same design. That is external validation produced by a human team, through an entirely different process, unaware that a comparison case existed — and subsequently endorsed by a regulator through breakthrough designation. The figure 37,000, by contrast, is not evidence; it is a description of compute and orchestration scale. Keeping the two apart is what stops a single convergence being read as the general validity of a method.

Discount this

The large majority of the system's predictions have no experimental validation; the Merck convergence is retrospective and n = 1, and cannot be read as a hit rate. "About 50% more likely to reach market" is a retrospective association over historical targets, not a prospective result. The primary write-up circulated as a bioRxiv preprint; we could not retrieve the full text of the Science version, so the figures here follow Stanford's own announcement and secondary coverage.

Safety & red-teaming NEJM AIRAG chatbot9/25

A patient-facing RAG chatbot handed over its backend configuration and 1,000 most recent patient conversations to browser developer tools

What

Alfredo Madrid-García and Miguel Rujas ran a two-stage test against a live (anonymised) patient-facing medical chatbot: first Claude Opus 4.6 for initial probing, then manual verification with Chrome developer tools, inspecting network traffic and stored data. What came out included the system prompt, the model and embedding configuration, retrieval parameters, backend endpoints, the API schema, knowledge-base content, and the 1,000 most recent complete patient–chatbot conversations — with no authentication required at any point, contradicting the service's own privacy claims. The root cause was sensitive configuration exposed through client-server communication rather than kept server-side. The case study reached the clinical community this week through NEJM AI's perspective piece "When the Chatbot Leaks", with the underlying report at arXiv:2605.00796 (2026-05-01) and circulation via the Duke AI Health roundup of 2026-09-25.

Why it matters

This is not an attack that needed a security specialist; the authors deliberately stress that no specialist skills were required. What an auditor can do with a browser, anyone can do. For any provider putting a RAG assistant on its website, its messaging account or a patient app, the report doubles as a checklist that can be run the same afternoon: is the system prompt or retrieval configuration visible client-side, and does the conversation endpoint require authentication? It also punctures a common ordering error in medical AI debate — we argue about diagnostic accuracy while plenty of deployments have not yet secured the far more basic property that patient conversations cannot be read by strangers.

Discount this

A single anonymised deployment, with the vendor unnamed, so outsiders cannot verify whether it has been patched, nor extrapolate a failure rate across comparable products. The NEJM AI piece is a perspective rather than original research, and we could not retrieve that page's full text, so the technical detail here follows the arXiv report.

Decision support I3LUNGNSCLC · immunotherapy9/13

I3LUNG: explainable AI lifted physicians' disease-control prediction from 57% to 65% — then weakened in external validation, and adding imaging and pathology did not help

What

Nature Medicine published the clinical-usability study from the I3LUNG consortium, covering 2,396 patients with advanced non-small cell lung cancer on immunotherapy across six international centres. Twenty physicians, across 200 assessments, used an explainable decision-support tool built on clinical and blood variables, and their disease-control prediction accuracy rose from 57% to 65%. The paper reports two negative findings alongside it: the model weakened in external validation, and the multimodal versions adding imaging or pathology data did not consistently reproduce the gains. Nature Medicine, 2026-09-13

Why it matters

An eight-percentage-point absolute gain, on a question that genuinely changes management — whether to continue immunotherapy — is a rare instance of clinically meaningful human-AI collaboration. But what makes the paper worth reading is that it puts two negative results in the same publication: multimodal is not automatically better, and external validation degrades. The dominant narrative of the past two years has been that more modalities mean more performance; I3LUNG uses the same patient cohort to say that on this task cheap clinical and blood variables already absorb most of the signal, and the expensive imaging and pathology modalities did not reliably add to it. For hospitals budgeting multimodal foundation-model programmes that is a very practical cost signal.

Discount this

The physician arm is a usability study with 20 clinicians and 200 assessments, not an outcomes trial — nobody has shown that eight points of prediction accuracy translate into survival or quality-of-life differences for patients. A 57% baseline is close to a coin flip, and a low starting point makes gains look larger. We could not retrieve the full text directly; the figures are cited from the September evidence briefing and its restatement.

Operational prediction JAMA Network Opendischarge date9/3

AI predicting the discharge date: a narrow win on admission day, and a widening loss as discharge approaches

What

A single-centre quality-improvement study compared an AI tool against case managers on discharge-date prediction across 22,349 inpatient encounters and 17,173 patients. Mean absolute error: 4.20 days for AI versus 4.27 for case managers at admission (a narrow AI win); 1.59 versus 1.29 at 48 hours before discharge (AI behind); 1.93 versus 0.98 at 24 hours (AI error close to double). JAMA Network Open, 2026-09-03

Why it matters

Discharge-date prediction is the backbone of bed management, throughput and discharge planning, and the shape of this result is instructive: AI wins at the point with least information and least actionable value, humans win at the point where information is sufficient and the prediction is actually used to schedule beds. The reason is not hard to guess — case managers hold exactly what the model cannot see: when family can collect, whether a long-term-care bed is free, how a given attending behaves. It also explains why so many "AI improves operational efficiency" projects fail to show benefit after go-live: the segment of error they improve was never the decision bottleneck. Used well, a tool like this belongs as a day-one rough estimate and resource flag, not as the basis for pre-discharge scheduling.

Discount this

Single centre, quality-improvement design, so it does not extrapolate across different case mixes and care processes; and the case managers' predictions may be partly self-fulfilling, since the same people also arrange the discharge. We could not retrieve the full text directly; the figures are cited from the September evidence briefing's restatement.

Imaging research Johns Hopkins · MESARadiology9/23

Opportunistic AI bone density on chest CT predicted later white-matter disease and cognitive decline — most clearly in people with diabetes

What

A Radiology study led from Johns Hopkins' radiology department ran a secondary analysis of the prospective MESA cohort (Multi-Ethnic Study of Atherosclerosis), using deep learning to measure vertebral bone mineral density (vBMD) opportunistically on non-contrast chest CT and matching it to later cognitive testing and brain MRI. Sample: 715 participants (CT plus cognitive testing) and 405 (brain MRI), median age 69. Findings: each 0.1 g/cm³ lower baseline vBMD corresponded to 0.025 standard deviations faster annual cognitive decline; white-matter hyperintensity in the corpus callosum grew 12.8% per year; fractional anisotropy in the anterior limb of the internal capsule fell 0.0048 SD per year; and participants with diabetes accumulated total white-matter hyperintensity roughly 5% faster. Diagnostic Imaging, 2026-09-23 (author interview: same outlet, 9/25)

Why it matters

Opportunistic screening is the cleanest business model in medical AI: the scan is already done, the radiation already taken, the fee already paid, and the marginal cost of AI is one model run for one extra risk-stratification variable. This study extends that logic across organ systems — a bone signal in the chest pointing at the brain's ageing trajectory — and if it holds, it means every chest CT already on file is an unread brain-health risk record. It is also one of the rare AI applications that creates value without persuading a clinician to change behaviour, which is a pointed contrast with the earlier items where nothing happens if physicians stop using the tool.

Discount this

This is association, not causation, and it is a secondary analysis of a cohort not designed for the question; the absolute effect sizes are small (0.025 SD per year) and the threshold for clinical use still needs prospective validation. The brain MRI subsample is only 405 people and the diabetes subgroup smaller still, and the confidence interval around that 5% difference is not given in the secondary coverage we cite. We could not retrieve the Radiology paper directly; the figures come from Diagnostic Imaging's report.

Evaluation methodology Microsoft ResearchGPT-5 · MedGemma · Claude6/26

Frontier models' health benchmark scores do not survive the simplest adversarial test — remove the critical image and they still answer correctly

What

Microsoft Research Health & Life Sciences (Yu Gu, Jingjing Fu, Xiaodong Liu and 29 further authors) published systematic adversarial stress tests in Nature Medicine against GPT-5, Gemini, Claude 3.5 Sonnet, GPT-4o, DeepSeek-VL2, Qwen3-VL, LLaVA-Med and MedGemma, across VQA-RAD, OmniMedVQA, MIMIC-CXR-VQA, PathVQA, SLAKE and MMMU. The core findings: models still guessed the right answer with critical visual information removed, minor prompt modifications degraded performance, and they produced convincing but flawed reasoning traces. The study also shows that popular health benchmarks vary widely in what they actually measure. Nature Medicine, 2026-06-26

Why it matters

"Still correct with the image removed" is unusually hard to argue with: it directly demonstrates that part of the score comes from statistical cues in the question rather than from reading the image. Regulatorily this lands at a pointed moment — the FDA has no settled AI benchmarking guidance while the market already carries many products whose principal clinical argument is a benchmark score. The conclusion is unflattering and useful: the gap between benchmark success and the evidence deployment requires is too wide for one to substitute for the other. The right test for a medical multimodal model is an ablation — remove what it is supposed to look at and see whether it still answers — not another pass at the leaderboard.

Discount this

This is benchmark-level work with no real clinical workflow involved, so it cannot be read as a safety verdict on any marketed device; and Microsoft is a commercial partner to some of the models tested, which belongs in the reading. The Claude version tested is 3.5 Sonnet, no longer the newest in that line, so cross-generation comparisons need care.

Evidence infrastructure UNC Health · Duke HealthJoint Commission · Hackensack9/22

The capacity to validate AI is itself becoming infrastructure: a statewide rural AI evaluation network in North Carolina, and the first Joint Commission AI certification in the US

What

UNC Health, Duke Health and North Carolina partners launched a statewide network to help rural and critical-access providers evaluate and implement AI tools, funded by The Duke Endowment at $4.4 million over three years (UNC, 2026-09-22). The same week, Hackensack Meridian Health received the first Joint Commission AI certification in the country, a review covering AI inventory, risk assessment, post-deployment monitoring and patient-data protections (TechTarget, collated by HIStalk, 2026-09-23).

Why it matters

The six studies above converge on one point: the value of medical AI depends heavily on local deployment conditions — workload, workflow, staff habits — and therefore cannot be delivered once and for all by a vendor's multicentre trial. That creates a structural problem: an academic medical centre can run its own DECIDE-AI-style prospective evaluation, a rural hospital cannot, and rural hospitals are exactly the buyers most exposed to vendor claims. The North Carolina network and the Joint Commission certification are two patches on the same gap: one supplies who does the validating, the other what counts as validated. Note that neither is the FDA — the governance layer is growing ahead of the written regulation.

Discount this

The Joint Commission's AI certification examines process and governance, not product efficacy; holding it does not mean any AI tool the organisation uses has evidence of clinical benefit. And $4.4 million over three years across a whole state's rural providers is small relative to the need, with no outcome metrics published so far.

02 — Product Analysis

One has prospective controlled evidence and no commercial path; the other has a beautiful AUC and no prospective evidence. Today's two ends are each missing the other's half

SHAKED

LLM clinical decision support for the emergency department · built in-house at Rambam Health Care Campus (Israel), with the Technion and Mayo Clinic

Function and position. In the emergency department, it proposes workup and consultation actions from the patient's current data. It is not a commercial product but a system built in-house in order to be evaluated rigorously. Its distinction is the evaluation design: two parallel emergency wings, one using it and one not, reported on 1,138 patients to DECIDE-AI stage 1 specification — a controlled condition almost no commercial vendor can obtain, because no hospital will cut its own emergency department in half to validate someone's software.

  • Strength : the safety numbers are solid — expert review rated 99 of 100 sampled outputs clinically appropriate with zero adverse events, obtained on real patients under real shift conditions rather than a retrospective dataset (Nature Medicine).
  • Concern : uptake falling 68% → 30% in four weeks (OR = 0.72 per shift hour) means any downstream benefit is diluted past detection — length of stay was 4.9 hours in both wings (P = 0.99) and the consultation cycle's −9.4 minutes did not reach significance (P = 0.077). State the missing data plainly: there is no comparison at all on diagnostic accuracy, readmission or mortality, and the authors themselves write that the evidence does not support deployment.
  • Transferable : the adoption curve is itself a deliverable. Any organisation deploying clinical AI should collect the regression of usage against shift workload as a required metric — it tells you whether the tool survives its first quarter earlier than any accuracy figure does.

PRISM2

Slide-level multimodal pathology foundation model · Tempus (US), data from Memorial Sloan Kettering and other institutions

Function and position. Trained on 2.3 million whole-slide images and 14 million question-answer pairs derived from 700,000 pathology reports, covering 685,507 specimens from 200,692 patients; 4.6 billion parameters, a Perceiver slide encoder feeding a Phi-3 Mini language model, with a deliberate split between base embeddings (for transfer learning) and diagnostic embeddings (for clinical tasks). Nature Medicine, 2026-07-31 (vol 32, 3203–3213)

  • Strength : pan-cancer detection at AUC 0.967 with diagnostic embeddings, against 0.947 for the previous PRISM and 0.931 for TITAN; prompt-based inference matched or exceeded clinical-grade products (P < 0.05) in prostate, breast and breast lymph node detection; and colorectal recurrence-free survival at C-index 0.809 beat a dedicated survival model's 0.773. Prompt-based inference means new tasks need no per-task labelling (Nature Medicine).
  • Concern : there is no prospective clinical trial at all — every "matched or exceeded clinical-grade products" comparison is retrospective, and item seven above, the Microsoft adversarial study, is precisely a warning against that reading: models answer correctly with critical information removed, so an AUC is not a deployment argument. A second concern is data concentration: the core training data comes from the MSK system, and distribution shift at external institutions is not quantified in the summary available to us.
  • Ask the vendor : the dual-track design of diagnostic and base embeddings is in practice an admission that one model cannot share a representation across tasks. Which track post-deployment monitoring watches, and how drift is declared, currently has no public answer.

03 — Companies & Competition

Who stands where, on what, against whom
Company / institution Recent state & numbers Position & moat
Tempus
Precision-medicine data and pathology foundation models
PRISM2 published in Nature Medicine on 2.3M WSIs, 14M QA pairs and 4.6B parameters, with pan-cancer AUC 0.967 and colorectal RFS C-index 0.809; co-authors span Tempus and Memorial Sloan Kettering (Nature Medicine, 2026-07-31). The moat is the scale at which slides, reports and clinical outcomes are paired, which academia can barely reproduce. The thin part is the absence of a prospective trial, plus a data lineage very close to that of Paige and PathAI — once MSK-grade data stops being exclusive, the advantage reverts to distribution.
Artera
FDA-cleared digital pathology risk stratification
ArteraAI Breast received FDA clearance as the first digital-pathology risk-stratification tool in breast cancer, for early-stage HR+/HER2− patients, outputting a distant-metastasis risk score with same-day results inside standard pathology workflow and no additional tissue; supporting data from the 2025 San Antonio Breast Cancer Symposium (ITN, 2026-05-06). The moat is an actual clearance embedded in existing workflow, requiring no change to tissue handling — which is exactly what item one implies: the tools that survive are the ones that do not ask clinicians to behave differently. The thin part is that prognostic is not predictive of treatment benefit, and guideline-level adoption still needs randomised evidence.
Stanford/James Zou 實驗室
Multi-agent drug discovery (academic, not a company)
The Virtual Biotech runs on roughly 37,000 agents, and its CD276 lung-cancer ADC design was independently arrived at by Merck and granted FDA breakthrough designation; the underlying Paperclip platform is said to cut time and cost by over an order of magnitude (Stanford Medicine, 2026-09-17). The moat is the orchestration method and AI-native data infrastructure rather than any single model — which also means large pharma can reproduce it internally. The thin part is that the vast majority of outputs have no wet-lab validation, and the only hard evidence so far is a single n = 1 external convergence.
Microsoft Research Health & Life Sciences
Evaluation methodology — in effect, the referee's chair
Adversarially stress-tested eight models including GPT-5, Gemini, Claude 3.5 Sonnet and MedGemma in Nature Medicine, showing that models still answer correctly with critical images removed and lose points on minor prompt changes (Nature Medicine, 2026-06-26). The moat is agenda-setting: before the FDA's AI benchmarking guidance settles, whoever defines valid evaluation indirectly defines who passes. The thin part is an obvious conflict of interest — it is also a commercial partner to some of the models tested.
Rambam · Technion · Mayo Clinic
In-house CDS and prospective evaluation capability
SHAKED completed a four-week DECIDE-AI stage 1 evaluation across 1,138 patients and two parallel emergency wings, and published the negative uptake result of 68% → 30% itself (Nature Medicine, 2026-08-19). The moat is something vendors cannot buy: the authority to split its own emergency department into a controlled comparison, and the freedom to publish a negative result without hurting a share price. The thin part is the absence of any commercial path — good evaluation methodology does not become a product by itself.
The Joint Commission
Healthcare accreditor, now issuing AI certification
Issued the first AI certification in the US to Hackensack Meridian Health, reviewing AI inventory, risk assessment, monitoring and patient-data protections (TechTarget / HIStalk, 2026-09-23). The moat is an existing accreditation channel: hospitals are already surveyed by it, AI is one more chapter, and adoption cost is far below that of new regulation. The thin part is the same fact — it reviews process, not efficacy, and is easily misread as an endorsement of quality.

Close it in one sentence: today's competition is not model against model but whoever can produce deployment-grade evidence against whoever can only produce an AUC. The right-hand column is crowded and capital-rich; the left-hand column currently holds little more than academic medical centres willing to split their own emergency departments in two, and they have no commercial incentive to turn the method into a product. That mismatch is the most structural gap in medical AI right now, and it is precisely why third-party validation infrastructure of the Joint Commission and North Carolina variety is starting to grow.

04 — Taiwan Angle

No local Taiwanese clinical-research news today, so this section works through what three international results mean once translated into Taiwan's institutions

(1) Taiwan's AI-SaMD registration pathway can demonstrate accuracy; it cannot demonstrate how many people are still using the tool in week three. The TFDA has stood up a dedicated smart medical device project office and runs an AI/ML medical device information and matchmaking platform, and Taipei Veterans General Hospital offers AI-SaMD clinical trial services, so the registration route is comparatively complete. But the decisive variable the Rambam trial identified — uptake collapsing with shift workload (OR = 0.72 per hour) — appears in no mandatory field of any registration dossier. For Taiwanese hospitals that means one more item in the deployment evaluation design: measure not only accuracy but actual usage at weeks 1, 4 and 12, and its regression against how busy the shift was. That number is cheap, collectable in-house, and predicts whether a purchase becomes a dormant licence better than any vendor's multicentre dataset.

(2) The chatbot leak report applies immediately to the messaging-account assistants Taiwanese hospitals run. What the researchers obtained was backend configuration plus the 1,000 most recent complete patient conversations, using nothing but browser developer tools and a commercial LLM, with no authentication at any point (arXiv:2605.00796). Many Taiwanese providers build patient-education and waiting-room Q&A on messaging platforms or a web front end, and what patients disclose in those conversations is history and symptoms — which under Taiwan's Personal Data Protection Act falls in the special category covering medical records and health examinations, with materially tighter handling limits than ordinary personal data. Two things can be done today: open your own assistant's network requests in developer tools and confirm the system prompt and retrieval configuration are not visible client-side; and confirm the conversation-history endpoint requires authentication and does not accept a guessable sequential parameter. This needs no security team, only someone willing to press F12.

(3) Opportunistic screening is the class of medical AI whose reimbursement value is easiest to argue in Taiwan, because the scan has already happened. The economics of the Johns Hopkins study hold unusually well here: the national health insurance system already generates large volumes of non-contrast chest CT each year, including low-dose lung cancer screening, and the Executive Yuan's policy statement on AI in medical imaging already treats imaging as a priority. The point is that the marginal cost is one model run — no extra scan, no extra radiation, no extra visit. It is one of the few settings where cost-effectiveness can be argued as a second clinical variable extracted from the same examination, and that is the easiest story to tell in a reimbursement negotiation. The limits are equally clear: the MESA cohort's composition differs from Taiwan's (715 participants, median age 69), effect sizes are small, and carrying the thresholds over directly would not be safe — rerunning the association locally is a precondition.

05 — Further Reading

Five pieces, chosen on one criterion: reading them changes what data you ask a vendor for
  1. Prospective evaluation of a large language model clinical decision support system in the emergency department — Nature Medicine (2026-08-19)

    The one to read in full today. The point is not the accuracy figures but the adoption curve from 68% to 30%, and how the authors write "physicians stopped using it" as the principal finding rather than a limitation.

  2. How a team of AIs discovered a promising lung-cancer drug — Nature (2026-09)

    A news feature, but it makes clear why Merck's independent convergence is the only part of the story that counts as evidence — a useful correction to the attention the number 37,000 absorbs.

  3. When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI — arXiv:2605.00796 (2026-05-01)

    Read it as a checklist, not a paper. Every exposure it lists maps to a check you can run against your own system this afternoon.

  4. Evaluating the robustness and readiness of large frontier models in health AI applications — Nature Medicine (2026-06-26)

    Until the FDA's benchmarking guidance lands, this is the most complete account of why a benchmark score is not a deployment argument. The ablation section can be lifted straight into a procurement specification.

  5. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study — The Lancet Gastroenterology & Hepatology

    This issue keeps returning to "nothing happens if physicians stop using it"; this paper is the other face of the same problem — whether continuous use erodes the skill exercised without AI. The trade-off is only visible with both.

06 — References

References
  1. Prospective evaluation of a large language model clinical decision support system in the emergency department. Nature Medicine, 2026-08-19. nature.com · PubMed 42618632
  2. Virtual biotech company puts thousands of AI scientist agents to work on drug discovery. Stanford Medicine News, 2026-09-17. med.stanford.edu
  3. How a team of AIs discovered a promising lung-cancer drug. Nature (news feature), 2026-09. nature.com
  4. The Virtual Biotech: A multi-agent AI framework for therapeutic discovery and development. Science, 2026. science.org · preprint: bioRxiv
  5. Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck. VentureBeat, 2026. venturebeat.com
  6. Virtual Biotech Company Puts 37,000 AI Agents to Work on Drug Discovery. Singularity Hub, 2026-09-18. singularityhub.com
  7. When the Chatbot Leaks: Securing Patient-Facing Medical AI in the Age of Dual-Use Large Language Models. NEJM AI, 2026. ai.nejm.org
  8. Madrid-García A, Rujas M. When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI. arXiv:2605.00796, 2026-05-01. arxiv.org
  9. AI Health Roundup — September 25, 2026. Duke AI Health, 2026-09-25. aihealth.duke.edu
  10. Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC (I3LUNG). Nature Medicine, 2026-09-13. nature.com
  11. Comparison of an AI tool and case managers for inpatient discharge-date prediction. JAMA Network Open, 2026-09-03. jamanetwork.com
  12. Healthcare AI News and Regulation: September 2026 Evidence Briefing. Veroscribe, 2026-09. veroscribe.com
  13. September 2026 healthcare AI briefing separates evidence from vendor announcements. Complete AI Training, 2026-09. completeaitraining.com
  14. Study: AI BMD assessment on chest CT linked to faster cognitive decline and white matter changes on brain MRI (Radiology). Diagnostic Imaging, 2026-09-23. diagnosticimaging.com
  15. New CT/MRI research may reveal links between bone loss, diabetes and cognitive decline (interview with Dr. Shadpour Demehri). Diagnostic Imaging, 2026-09-25. diagnosticimaging.com
  16. Advances in AI — September 2026. Diagnostic Imaging, 2026-09. diagnosticimaging.com
  17. Gu Y, Fu J, Liu X, et al. Evaluating the robustness and readiness of large frontier models in health AI applications. Nature Medicine, 2026-06-26. nature.com
  18. Vorontsov E, Shaikovski G, Liu S, et al. End-to-end multimodal pathology foundation model with clinical dialogue (PRISM2). Nature Medicine 32:3203–3213, 2026-07-31. nature.com
  19. FDA Clears AI Digital Pathology Tool in Breast Cancer (ArteraAI Breast). Imaging Technology News, 2026-05-06. itnonline.com
  20. Statewide network launches to help rural and critical access providers harness AI for better patient care. UNC, 2026-09-22. unc.edu
  21. Inside Hackensack Meridian's first-in-nation Joint Commission AI certification. TechTarget, 2026-09. techtarget.com
  22. Healthcare AI News 9/23/26. HIStalk, 2026-09-23. histalk2.com
  23. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology. thelancet.com
  24. 食品藥物管理署智慧醫療器材專案辦公室成立. 衛生福利部. mohw.gov.tw
  25. 智慧醫療器材資訊暨媒合平台. 衛生福利部食品藥物管理署. aimd.fda.gov.tw
  26. AI-SaMD 產品臨床試驗服務. 臺北榮民總醫院人工智慧中心. vghtpe.gov.tw
  27. AI 在醫療影像之應用(院會議案). 行政院. ey.gov.tw
Editor's note: (1) Secondary citation: we could not retrieve the primary pages for I3LUNG (Nature Medicine, 9/13) or the discharge-prediction study (JAMA Network Open, 9/3), so all figures come from the September evidence briefings listed as references 12 and 13; the Radiology bone-density and brain study is likewise cited from Diagnostic Imaging's coverage rather than the paper. (2) Access limits: the NEJM AI perspective "When the Chatbot Leaks" and the Science Virtual Biotech paper could not be retrieved in full in this production environment; technical detail follows the arXiv:2605.00796 report and Stanford's own announcement plus secondary coverage respectively, with links listed as found for verification. (3) Window: several primary papers here carry publication dates before this week (8/19, 7/31, 6/26) because they entered clinical discussion during 22–28 September via the Duke AI Health roundup of 9/25 and the September evidence briefings; dates are given as published and have not been shifted forward. (4) Thin local news: there was no Taiwanese clinical-research news of sufficient weight in the 22–28 September window, so section 04 is written as three institutional inferences rather than local reporting, and says so in its subhead. (5) Unaudited self-reported figures: Paperclip's "over an order of magnitude" cost reduction, PRISM2's comparisons against clinical-grade products, and ArteraAI Breast's SABCS supporting data are all reported by the research team or vendor and not independently audited. (6) This report is not investment or medical advice.