◆ AI & Medical AI Daily
–
Monday · Clinical Applications & Research

Clinical AI moved out of the clinic and into the operating room this week: CRISP spoke to 92.6% of intraoperative decisions off 100,000 frozen sections — while the one paper that actually put physicians in the loop moved their accuracy from 57% to 65%

Between Thursday and Sunday, Nature Medicine and npj Digital Medicine put out at least eight clinical-AI papers, running from intraoperative frozen sections to immunotherapy response prediction, an AI-native eye clinic, and rare-disease mining inside electronic health records. Laid side by side, two things are hard to miss. First, most of the corresponding institutions are in China — the intraoperative pathology foundation model CRISP (Sun Yat-sen University Cancer Center, Sep 10), the AI-native eye clinic AI-TEC (Beijing Tsinghua Changgung, Sep 10), HE2FISH, reading lymphoma gene rearrangements off H&E slides (CAMS Cancer Hospital, Sep 12), and the thyroid nodal-metastasis model Thy-CLNM (Hangzhou TCM Hospital, Sep 12) all come from there. Second, the only paper that actually brought physicians into the room and measured what human-plus-AI is worth is the EU-funded I3LUNG (Sep 13) — and its number is accuracy moving from 57% to 65%. Model scale and physician evidence grew in different places this week.

01 — Top Stories

Eight papers, two kinds of evidence: one side scaling models into workflows, the other measuring whether doctors actually got better
Intraoperative pathology Sun Yat-sen University Cancer CenterCRISP9/10

CRISP: an intraoperative pathology foundation model trained on 100,000 frozen sections that directly informed the surgical decision in 92.6% of prospective cases

What

On 10 September, a team at Sun Yat-sen University Cancer Center published CRISP in Nature Medicine — a pathology foundation model built specifically for intraoperative frozen sections. It was trained on more than 100,000 frozen sections from ten medical centers and evaluated on over 15,000 intraoperative slides across nearly 100 retrospective diagnostic tasks, with generalization tested across 6 institutions, 14 tumor types and 24 anatomical sites. The prospective figures: the model's output directly informed the surgical decision in 92.6% of cases, human–AI collaboration cut diagnostic workload by 35%, 105 ancillary tests were avoided, and micrometastasis detection reached 87.5% accuracy.

Why it matters

Frozen sections are one of the few places in a hospital where an AI that is one second late is worth nothing: the patient is on the table and the surgeon needs a yes or no inside twenty minutes — cut wider, or clear the nodes. Most pathology AI of the past few years has grown on digitized formalin-fixed slides, where image quality is good and time pressure is low; moving a model onto frozen, unevenly stained, fragmented intraoperative tissue is a different order of difficulty. CRISP puts prospective deployment, cross-institution generalization and a measured workload drop into one paper, which is still uncommon for pathology foundation models.

Discount this

"Directly informed the surgical decision in 92.6% of cases" measures how often the model's output entered the decision, not how often it made the decision more correct; at abstract level there is no head-to-head accuracy comparison against individual pathologists, only a claim of superiority over existing foundation models. The 35% workload reduction and the 105 avoided ancillary tests were both measured inside that center's own workflow.

Explainable AI I3LUNGNSCLC9/13

I3LUNG: with twenty physicians in the loop, explainable AI lifted immunotherapy-response accuracy from 57% to 65% — the week's only genuine human-plus-AI number

What

The EU Horizon Europe–funded I3LUNG project (led by IRCCS Istituto Nazionale dei Tumori in Milan) published its clinical-usability study of a multimodal explainable AI decision-support tool in Nature Medicine on 13 September. Six centers, 2,396 patients with advanced non-small-cell lung cancer, 339 of them with complete multimodal data. The models predict 6- and 24-month overall survival, disease control rate and overall response rate: AUC 0.74 for OS24 against 0.53 for PD-L1 in the same test set, and a C-index of 0.65 versus 0.55 for the composite LIPI score. Then comes the part that matters — 20 physicians, 10 experts and 10 non-experts, actually used it: sensitivity for disease control rose from 0.72 to 0.87, accuracy from 57% to 65%, and physicians adopted correct XAI suggestions 74.5% of the time.

Why it matters

Those eight percentage points, 57% to 65%, are the most honest number of the week. They say two things at once: explainable AI does help — and experts and non-experts benefited similarly, which suggests it is patching systematic judgment bias rather than a knowledge gap — and even with an AUC-0.74 model at their elbow, physicians still land at 65%. The loss between a model's discrimination and the quality of the clinical decision is real. Nature Medicine's accompanying commentary, A multimodal murmuration for immunotherapy, takes up the same multimodal-integration thread.

Discount this

Retrospective design on heterogeneous real-world data; only 339 patients had complete multimodal records, and the multimodal gains seen in training did not consistently replicate in test and validation sets; a single external validation cohort with different baseline characteristics; sparse genomic data, radiomic features limited to the primary lesion, and roughly 15% of CT scans not acquired at the start of immunotherapy. The 20-physician usability test is also a small sample.

AI-native care AI-TECBeijing Tsinghua Changgung9/10

AI-TEC: an early reckoning from an "AI-native" eye clinic in China — from assistive tool to the care model itself

What

On 10 September Nature Medicine ran "Initial lessons from real-world implementation of an AI-agent eye clinic in China", from the Beijing Visual Science and Translational Eye Research Institute (BERI) at Beijing Tsinghua Changgung Hospital, Tsinghua University. The piece is not about one model's AUC but about what was learned from rebuilding a clinic around AI agents as the spine of care: success turns on workflow integration, clinician engagement and measurable clinical value, and the authors note that expert-reviewed, high-quality data improves AI-TEC's performance.

Why it matters

"AI-assisted" and "AI-native" carry completely different institutional costs. The first hangs a model beside an existing pathway and keeps a human in the loop when it errs; the second rebuilds the pathway around the agent, which means liability, scheduling, referral and escalation all have to be rewritten. Ophthalmology was pushed there first — single imaging modality, standardized reading, high patient volume — so the potholes it hits are likely the ones other specialties will hit. The value here is that it is a retrospective on an implementation, not a pre-deployment vision.

Discount this

This is a commentary and implementation review, not a controlled trial. The full text sits behind a paywall; the publicly visible portion carries no patient counts, no concordance rate against human clinicians and no safety-event statistics. This report is written from the abstract and visible sections only, and does not extrapolate figures that were not published.

Digital pathology HE2FISHCAMS9/12

HE2FISH: predicting MYC/BCL2/BCL6 rearrangements from ordinary H&E slides, external AUC above 0.81, stratifying survival comparably to FISH

What

On 12 September a team at the National Cancer Center / Cancer Hospital, Chinese Academy of Medical Sciences published HE2FISH in npj Digital Medicine: predicting MYC, BCL2 and BCL6 rearrangement status — and co-rearrangement patterns — in diffuse large B-cell lymphoma straight from routine H&E sections. The data covers 1,377 patients with paired H&E and FISH across five hospitals; mean external AUC exceeds 0.81 for single-gene prediction, and in survival analysis HE2FISH-predicted rearrangement status stratified overall and disease-free survival comparably to FISH.

Why it matters

FISH is expensive, slow and tissue-hungry, and the double-hit / triple-hit classification directly decides whether first-line chemotherapy is intensified. If an H&E slide that had to be stained anyway can screen out most patients who never needed FISH, the saving is not only the assay cost but the decision time. This is the most practical commercial line pathology AI currently has: not replacing the confirmatory test, but triaging in front of it.

Discount this

The framework is scoped to DLBCL cases with diffuse growth patterns, and the authors note that adding structured clinicopathological data was needed to lift performance in specific applications — meaning the image-only signal has a ceiling. 1,377 paired cases across five hospitals is large for lymphoma, but once stratified by the rarer combinations across three genes the per-subgroup counts stay thin.

Rare disease RDMAUIUC9/10

RDMA: mining rare diseases out of clinical notes with no task-specific training and a small quantized model — 10× cheaper inference, 17× lighter hardware

What

On 10 September, computer scientists at the University of Illinois Urbana-Champaign, with collaborators at the University of Illinois College of Medicine Peoria, published RDMA in npj Digital Medicine: an agent framework with built-in abbreviation resolution, implicit phenotype reasoning and ontology grounding against Orphanet and HPO, used to surface rare-disease mentions in unstructured clinical text. The paper claims it beats fine-tuned and RAG baselines without task-specific training, cuts inference cost by up to 10× and local hardware requirements by up to 17×, and adds uncertainty flagging to reduce expert annotation burden. Code is public on GitHub.

Why it matters

The paper notes that roughly one in ten Americans lives with a rare disease, yet over half of Orphanet codes have no direct ICD equivalent — meaning coding-based epidemiology is structurally blind to them. What is interesting about RDMA is not the accuracy but the claim it stakes on small quantized models running locally: for a hospital, keeping notes inside the building and running on ordinary hardware decides adoption far more than a few points of F1.

Real-world evidence RespondHealthGLP-19/10

Turning EHRs into computable patient journeys: language models extract entities, medical ontologies bind them, and GLP-1 response is the first demonstration

What

On 10 September Nature Medicine published a method for building longitudinal patient journeys out of structured and unstructured EHR data, first-authored from RespondHealth in Bethesda, Maryland, with collaborators at Stanford, the University of Miami, Drexel, Penn and Mount Sinai. Large pre-trained language models extract clinical entities from free text, medical ontologies organize them into knowledge graphs, and the result is integrated with structured data into patient-level trajectories. The paper reports physician adjudication against a blinded expert reference standard showing high accuracy and high inter-reviewer agreement; the worked example is GLP-1 receptor agonist therapy — identifying treatment initiators, modeling weight and HbA1c change, and characterizing real-world response.

Why it matters

What has held real-world evidence back is not statistical method but the fact that most of what is in a chart is prose. If language models can turn free text into computable entities reliably, the usable sample for RWE studies expands at a stroke — with more consequence for reimbursement negotiation, post-market surveillance and the kind of regulatory questions this daily covers on Wednesdays than any one more diagnostic model. The authors are careful to say human oversight remains essential and that design and interpretation stay with investigators; nothing here is claimed as fully automated.

Discount this

The publicly readable portion gives no patient or note counts, only "large-scale, across all clinical conditions"; accuracy is described as high without a stated figure. The first-author institution is a commercial company and the demonstration topic is GLP-1, precisely where pharma demand is heaviest — worth holding the incentive structure in mind while reading.

Digital twins ICDT-HM9/11

21 authors across China, the UK, the US, France and Singapore launch an international consortium for medical digital twins — a blueprint, with no quantified commitments yet

What

On 11 September Nature Medicine published "A global digital navigator of human health for precision medicine", announcing the International Consortium of Digital Twins in Healthcare and Medicine (ICDT-HM) — 21 authors from universities and medical centers in China, the UK, the US, France and Singapore. The aim is to establish medical digital twins as new infrastructure for precision health; the paper sets out a trajectory-inference framework for an individual digital twin and a roadmap "from blueprint to a global digital twin ecosystem."

Discount this

This is a declaratory document, not a research result. There are no numerical targets, no binding commitments and no funding figures — only a phased roadmap. It reads more usefully as a signal of who intends to align with whom than as a technical advance.

Preoperative decision Thy-CLNM Fusion9/12

Thy-CLNM: 934 papillary thyroid carcinoma cases, ultrasound habitat radiomics plus deep learning for central lymph node metastasis, validation AUC 0.862

What

On 12 September a team at Hangzhou TCM Hospital, affiliated with Zhejiang Chinese Medical University, published the Thy-CLNM Fusion model in npj Digital Medicine, predicting central lymph node metastasis in papillary thyroid carcinoma from ultrasound using deep learning combined with habitat radiomics, in 934 patients who underwent cervical nodal surgery between January 2023 and December 2024. Validation-cohort AUC was 0.862, sensitivity 78.9%, specificity 79.3%. Low-risk cases activated in hyperechoic regions while high-risk cases localized to hypoechoic areas associated with aggressive pathology. The study is retrospective and registered at ClinicalTrials.gov as NCT06725628.

Why it matters

Whether to perform prophylactic central-compartment neck dissection in papillary thyroid carcinoma is a question surgery has argued over for years — over-dissect and you risk the parathyroids and recurrent laryngeal nerve, skip it and you may miss metastasis. An AUC of 0.862 is not enough to replace intraoperative judgment, but it moves the decision from experience alone toward a quantifiable prior, using ultrasound that already exists preoperatively and adding no new test.

Discount this

Single center, retrospective, a two-year window; sensitivity and specificity both sit around 79%, meaning roughly one patient in five lands on the wrong side. Ultrasound is heavily operator- and machine-dependent, and cross-institution generalization has not been tested.

02 — Product Analysis

CRISP and I3LUNG: one bets on workflow embedding, the other on physician evidence — and the two have not met yet

CRISP

Intraoperative frozen-section pathology foundation model · Sun Yat-sen University Cancer Center (China)

Function and position. It moves the pathology foundation model out of the formalin lab and into the operating room, aimed at high-volume surgical centers with intraoperative pathology demand and thin pathologist staffing. Trained on 100,000-plus frozen sections from ten centers; evaluated across 6 institutions, 14 tumor types and 24 anatomical sites (Nature Medicine, Sep 10).

  • Strength : scale and generalization arrived together. In prospective testing the model's output entered the surgical decision in 92.6% of cases, workload fell 35% and 105 ancillary tests were avoided — rare for a pathology foundation model to report deployment numbers at all (source).
  • Strength : 87.5% accuracy on micrometastasis detection lands precisely on what frozen sections most often miss and what most changes the extent of dissection (source).
  • Concern : there is no arm comparing patient outcomes with and without CRISP. 92.6% is an adoption rate, not a correctness rate, and no head-to-head figure against pathologists appears at abstract level. What is missing is missing; the paper does not supply it and neither will this report.
  • Concern : frozen-section preparation quality varies enormously between hospitals, and the abstract does not say whether the ten training centers span the quality range a smaller hospital would produce.

I3LUNG

Multimodal explainable AI for immunotherapy decisions · led by IRCCS Istituto Nazionale dei Tumori, funded by EU Horizon Europe (Italy)

Function and position. It fuses clinical, radiomic and genomic features to predict immunotherapy survival and response in advanced NSCLC, and exposes the reasoning to oncologists through an explainable interface. 2,396 patients across six centers (Nature Medicine, Sep 13); project background at the CORDIS project page.

  • Strength : it puts humans in the experiment. Twenty physicians (10 experts, 10 non-experts) used it — DCR sensitivity 0.72→0.87, accuracy 57%→65%, correct XAI suggestions adopted 74.5% of the time. That remains uncommon in clinical AI papers (source).
  • Strength : the comparators are honest. Not measured against nothing, but against PD-L1 (AUC 0.53), ECOG PS, NLR, LDH and LIPI — the markers oncologists actually use (source).
  • Concern : the multimodal selling point did not hold in the test set — gains seen in training failed to replicate consistently, and only 339 patients had complete multimodal data. A system sold as multimodal whose strongest result comes from the clinical-feature model owes an explanation.
  • Concern : 65% accuracy is a long way from clinically acceptable. Deployed as a formal decision tool, a third of the judgments are still wrong; the defensible positioning today is discussion aid, not decision authority.

03 — Companies & Competition

Who stands where, on what, against whom
Institution Recent state & numbers Position & moat
Sun Yat-sen University Cancer Center
Guangzhou · CRISP
Published CRISP on Sep 10: trained on 100,000-plus frozen sections from ten centers, evaluated on 15,000+ intraoperative slides, 92.6% of prospective cases informing the surgical decision, workload down 35% (Nature Medicine). The moat is the frozen-section corpus itself — it accrues only from surgical volume and cannot be bought or synthesized. The weakness is the absence of outcome-anchored comparative evidence; the moment a regulator asks for efficacy endpoints rather than adoption rates, what exists is not enough.
IRCCS Istituto Nazionale dei Tumori / I3LUNG
Milan · EU Horizon Europe consortium
Published its multimodal XAI usability study on Sep 13: 2,396 NSCLC patients, 20 physicians in the loop, accuracy 57%→65%, OS24 AUC 0.74 against PD-L1's 0.53 (Nature Medicine; project details at CORDIS). The moat is methodological credibility: an EU-funded multinational consortium willing to publish an unflattering eight-point human–AI gain, which is an asset in regulatory and reimbursement conversations. The weakness is that neither its model performance nor its scale leads.
Beijing Tsinghua Changgung / BERI
Beijing · AI-TEC eye clinic
Published an implementation review of its AI-native eye clinic on Sep 10, naming workflow integration, clinician engagement and measurable clinical value as the three conditions (Nature Medicine). Patient counts and concordance rates are not in the publicly visible sections. The moat is institutional permission — being allowed to rebuild an entire clinic around agents is a governance question, not a technical one, in most health systems. The weakness is the same fact: that permission does not export to other regulatory environments.
Chinese Academy of Medical Sciences 腫瘤醫院
Beijing · HE2FISH
Published HE2FISH on Sep 12: 1,377 DLBCL patients with paired H&E and FISH across five hospitals, mean external AUC above 0.81 for single genes, survival stratification comparable to FISH (npj Digital Medicine). The moat is paired data: cases with both H&E and FISH accumulate only at large cancer centers where molecular testing is routine. The competitors are commercial pathology-AI firms such as Paige and PathAI — they hold the regulatory pathway, the academic groups hold the data.
RespondHealth
Bethesda, MD · computable patient journeys
On Sep 10, with Stanford, Miami, Drexel, Penn and Mount Sinai, published a method using language models to extract EHR entities and assemble longitudinal journeys, demonstrated on GLP-1 response (Nature Medicine). The moat sits where academic endorsement meets pharma demand; the buyers are drug makers and payers who need real-world evidence. The competitors are incumbent RWE suppliers such as Tempus and Flatiron — the difference being a method published in a top journal rather than a white paper.
University of Illinois Urbana-Champaign
RDMA · open-source rare-disease mining agent
Published RDMA on Sep 10, claiming it beats fine-tuned and RAG baselines with no task-specific training, at up to 10× lower inference cost and 17× lighter hardware, with code on GitHub (npj Digital Medicine). The moat is open source plus local deployment: keeping notes inside the building persuades a hospital's security office more than any accuracy figure. The weakness is the absence of commercial support — maintenance lands on the hospital.

Today's competitive structure is two layers that do not touch. On the data-and-workflow layer, advantage concentrates where surgical and patient volume are largest — frozen sections, paired H&E and FISH, an AI-native clinic; none of this has a shortcut, it accrues from years of clinical throughput, and four of this week's papers landed in one region. On the methodology-and-evidence layer, advantage sits with the European and American consortia and academic groups willing to publish unflattering numbers: I3LUNG's 57%→65%, RDMA's cost-rather-than-accuracy claim — both written for the regulatory and procurement gate. The real problem is that the two layers have not met: the largest models lack outcome-anchored evidence, and the best-evidenced models lack scale. Whoever stacks them first sets the bar for the next round.

04 — Taiwan Angle

Taiwan's three centers map onto exactly what this week's papers are missing

(1) "Implementation, validation, reimbursement" is, today, precisely CRISP's three gaps. Taiwan's Ministry of Health and Welfare set up three smart-healthcare centers — a Center for Responsible AI in Healthcare spanning 8 hospitals across northern, central and southern Taiwan; a Center for External AI Validation with 4 hospitals, linking cross-institution data over FHIR and building a federated-learning platform; and a Center for Clinical AI Impact Evaluation with 5 hospitals, collecting cross-hospital clinical benefit and cost data for health-economic estimation and insurance pricing — explicitly framed around solving implementation, validation and reimbursement. Put this week's papers in that frame and it reads clearly: CRISP has implementation (92.6% of decisions) but no outcome-anchored validation material, and has not touched reimbursement; I3LUNG has the validation method (20 physicians in the loop) but no scale. Taiwan's move is not to train a bigger frozen-section model — it cannot out-volume Guangzhou — but to make the third center's work exportable: a cross-institution method for measuring clinical benefit and cost is exactly what this week's papers commonly lack.

(2) The EHR interoperability track is the precondition for this week's RespondHealth paper. At a Kaohsiung Medical University forum in late June, Executive Yuan Political Commissar Chen Shih-chung and Health Minister Shih Chung-liang laid out the "333 policy" — breaking down the walls between hospital information systems and standardizing data structures — alongside President Lai's deepening-health plan, budgeted at NT$48.9 billion over five years, and the commitment that medical centers' electronic health records will interoperate by the end of 2026, extending to regional and district hospitals thereafter. The computable-patient-journey method in Nature Medicine works only because there are large volumes of cross-institution longitudinal records to extract from. If Taiwan genuinely links its medical centers by year-end, it will hold both a single-payer's complete care trajectories and National Health Insurance cost data — a combination no US commercial RWE supplier can assemble. What remains is not technical but governance: who may use it, how far, and how patients consent.

(3) One honest line: Taiwan had no comparable clinical-AI publication this week. Within the range we surveyed, we found no Taiwan-led clinical AI paper in a journal of this tier published between 8 and 14 September. Last year's Responsible AI Center results showcase, in which ten hospitals built a trustworthy governance mechanism together, points the right way — but institution-building does not appear on Nature Medicine's weekly list, and that gap is worth sitting with rather than routing around.

05 — Further Reading

Chosen to put today's eight papers back in context: one on measurement, one on evidence, one on an RCT that failed, one from the floor
  1. Medical AI has a measurement problem — Nature (2026-07-28)

    Arjun Manrai's argument is that evidence-based medicine rests on a reliable yardstick, and medical AI's complexity has broken the yardstick itself. CRISP's "92.6% of decisions" is exactly the phenomenon described: measuring what is easy to measure rather than what matters.

  2. Show us the evidence for the value of medical AI — Nature Medicine (2026 年 4 月號)

    Nature Medicine's own editorial said it back in April: show the evidence for the value. Six months later, of the eight papers it and its sister journal ran this week, exactly one actually measured physicians. Reading them against each other is instructive.

  3. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial — Nature Medicine (2026-06-26)

    The most important negative result of the year: 16 primary-care facilities in Kenya, 103 clinical officers, 9,691 patients, a GPT-4o-based AI Consult 2.0 embedded in the EMR — 14-day treatment failure 2.2% versus 2.0%, adjusted OR 0.77 (95% CI 0.55–1.08, P=0.13), no significant difference. Read it after today's eight and you will treat AUC figures more carefully.

  4. Can AI fix health care? In the chaos of emergency rooms, the technology comes up short — STAT News (2026-09-09)

    Field reporting from the same week, pointing the opposite way to every paper here: models keep getting stronger in journals, while in the emergency department the time saved never leaves the building. Paywalled; cited here only from the publicly visible headline and standfirst.

  5. Safety, efficacy and acceptability of human-GenAI single-session exposure-based intervention for academic anxiety: randomized controlled trials — npj Digital Medicine (2026-09-12)

    Same week, easily missed: Tsinghua's Department of Psychological and Cognitive Sciences, 425 participants across two RCTs, a hybrid human–GenAI single-session exposure program for academic anxiety with effect sizes of d=0.25–0.35 against an active control. Read it for the design — it is the week's only randomized trial with a genuine active comparator.

06 — References

References
  1. A clinically-oriented foundation model for intraoperative pathology. Nature Medicine, 2026-09-10. nature.com
  2. Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine, 2026-09-13. nature.com
  3. A multimodal murmuration for immunotherapy. Nature Medicine, 2026-09-13. nature.com
  4. Initial lessons from real-world implementation of an AI-agent eye clinic in China. Nature Medicine, 2026-09-10. nature.com
  5. Computable longitudinal patient journeys from structured and unstructured EHR data. Nature Medicine, 2026-09-10. nature.com
  6. A global digital navigator of human health for precision medicine. Nature Medicine, 2026-09-11. nature.com
  7. Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma. npj Digital Medicine, 2026-09-12. nature.com
  8. Deep learning combined habitat radiomics analysis of central lymph node metastasis in papillary thyroid carcinoma. npj Digital Medicine, 2026-09-12. nature.com
  9. RDMA: cost effective agent-driven rare disease mining from electronic health records. npj Digital Medicine, 2026-09-10. nature.com
  10. RDMA source code repository. GitHub. github.com
  11. Safety, efficacy and acceptability of human-GenAI single-session exposure-based intervention for academic anxiety: randomized controlled trials. npj Digital Medicine, 2026-09-12. nature.com
  12. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine, 2026-06-26. nature.com
  13. Medical AI has a measurement problem. Nature, 2026-07-28. nature.com
  14. Show us the evidence for the value of medical AI. Nature Medicine, Volume 32, April 2026. nature.com
  15. Can AI fix health care? In the chaos of emergency rooms, the technology comes up short. STAT News, 2026-09-09. statnews.com
  16. U.K. unveils recommendations for regulating AI in medicine. STAT News, 2026-09-09. statnews.com
  17. I3LUNG — Integrative science, Intelligent data platform for Individualized Lung cancer care with Immunotherapy. IRCCS Istituto Nazionale dei Tumori. istitutotumori.mi.it
  18. I3LUNG project record, grant agreement 101057695. CORDIS, European Commission. cordis.europa.eu
  19. 臺灣智慧醫療三大中心. 衛生福利部. aicenter.mohw.gov.tw
  20. 高醫大論壇揭示AI醫療新局!衛福部推「333政策」 國家489億預算力挺. 聯合新聞網, 2026-06-27. udn.com
  21. 醫療AI要聰明,還要能負責!十家醫院打造可信賴的智慧醫療守門人:衛生福利部負責任AI執行中心成果發表會. 衛生福利部. mohw.gov.tw
Editor's note: (1) All eight top stories rest on journal abstracts and publicly visible sections; several Nature Medicine and npj Digital Medicine full texts sit behind a paywall, and any figure the abstract did not disclose — AI-TEC's patient counts and human concordance, the RespondHealth paper's note counts and stated accuracy — is marked as undisclosed rather than estimated. (2) Both STAT News items are paywalled and are cited only from the publicly visible headline and standfirst. (3) The observation that this week's clinical-AI papers skew toward Chinese institutions comes from our own manual count of the 8–14 September listings at Nature Medicine and npj Digital Medicine, not a census of all journals; treat it as a sample observation, not a statistical conclusion. (4) CRISP's "directly informed the surgical decision in 92.6% of cases" is an adoption metric as reported by the authors, not diagnostic accuracy, and is flagged as such in the text. (5) The three-center composition and the 333 policy in the Taiwan section come from the Ministry of Health and Welfare's site and a 27 June 2026 UDN report respectively — not this week's news, used as institutional background. (6) Today's rotation is clinical applications and research, so regulatory and funding items from the same week — the UK's AI-in-medicine regulatory recommendations of 9 September, ARPA-H's heart-failure autonomous-AI investment of 9 September — are not among the top stories; the former is listed in the references for the record.