92.7% on text, 76.9% the moment there is a picture — and this week's hardest evidence says the fastest thing to fall out of the note is the patient
This week's clinical evidence splits along one seam. A 1,477-question haematology benchmark posted to medRxiv on September 2 by a Dresden team put Claude Opus 5 at 92.7% on text-only questions and 76.9% on the multimodal ones; Gemini-3.1 Pro scored 91.4% against 78.7%. A second preprint posted the same day found frontier LLMs triaging vignettes at 73.3–86.7% against 48.9% for Healthdirect, Australia's government-backed symptom checker. And on September 3 a University of Edinburgh review in BMJ Digital Health & AI changed the yardstick from minutes saved to what goes missing from the note: expression, gesture, affect, and whatever the patient stops saying once they know the room is being recorded.
01 — Top Stories
Edinburgh review: ambient AI scribes drop the patient's own account — and 40% of UK GPs already use one
On September 3, Lucas Seuren, Robin Williams and Kathrin Cresswell of the University of Edinburgh published a NASSS-framed review in BMJ Digital Health & AI (vol 2, no 1; DOI 10.1136/bmjdh-2026-000091), searching Embase, PubMed and Scopus to March 20, 2026 and including 27 articles, 13 of them empirical. The university's September 4 release notes that roughly 40% of UK general practitioners now use an ambient scribe.
Nearly every scribe study to date has measured minutes saved. This review changes the instrument: a clinical note is co-produced in the encounter, and the model only hears the audio. Facial expression, gesture and affect never enter the text stream; patients pull back on substance use, domestic violence and mental health once they know a recorder is running; and writing is itself a cognitive act, so handing it over is cognitive offloading — clinicians in the reviewed literature reported not recognising their own notes and not remembering the patient at follow-up. The authors argue ambient scribes must be evaluated as complex interventions, not productivity tools, with contextual validation and longitudinal follow-up.
This is a framework-driven review of 27 papers, not new prospective data, and it pools no effect sizes; the 13 empirical studies vary widely in setting and sample. The 40% figure comes from an external survey relayed in the press release, not from the review's own measurement.
1,477 haematology questions: ten frontier models approach 93% on text and lose 13–16 points the moment an image is attached
Radoynova, Benouis, Middeke, Eckardt and colleagues at University Hospital Carl Gustav Carus (TUD Dresden), the Else Kröner Fresenius Center for Digital Health and NCT/UCC Dresden posted a ten-model benchmark to medRxiv on September 2, covering 1,477 board-style multiple-choice questions drawn from five teaching banks run by the British Society for Haematology, the American Society of Hematology and EBMT, spanning nine disease areas and six clinical skill domains, in text-only and multimodal form. Claude Opus 5 scored 92.7% text and 76.9% multimodal; Gemini-3.1 Pro 91.4% and 78.7%; Gemini-3.6 Flash 91.0% and 74.8%; GPT-5.6 Sol 89.9% and 76.7%.
Half of haematology's diagnostic entities live in pictures — peripheral smears, marrow aspirates, flow plots. A model that reaches 93% on prose retains three-quarters of that on the images, which is where the specialty actually works. The authors log two further findings: accuracy correlates significantly with model size, but this generation's gains landed almost entirely on open-weight models while proprietary ones improved only marginally; and, more importantly, the top performers show highly concordant failure patterns — they share the same blind spots on hard cases. That undercuts the most common risk mitigation in clinical practice, which is asking a second model. The authors conclude that continuous expert-on-the-loop monitoring remains non-negotiable.
A preprint, not peer reviewed. Multiple choice is not clinical decision-making: the options are supplied, the history is complete, nothing is time-pressured and no result is equivocal. These teaching banks have been public for years and plausibly sit inside the training data, so some of the text score may be recall rather than reasoning. Per-model figures are taken from the abstract page; we did not obtain the full supplementary tables.
Frontier LLMs triage at 73–87%, Australia's government-backed symptom checker at 48.9% — and the critical-miss count is 0 to 2
Azwad Raza Chowdhury (Victorian Institute of Technology, Melbourne) and Bushra Chowdhury (CQUniversity Australia) posted a standardised vignette evaluation to medRxiv on September 2: 45 Semigran vignettes (15 emergency, 15 non-emergent, 15 self-care) run against Australia's government-backed Healthdirect and six LLM configurations across ChatGPT, Claude and Gemini at free and paid tiers. Healthdirect triaged at 48.9% (95% CI 35.0–63.0) with Cohen's kappa 0.233, emergency sensitivity 46.7% and two critical misses. The LLMs ranged 73.3–86.7%, kappa 0.600–0.800, emergency sensitivity 80.0–86.7%, with zero critical misses across 270 evaluations (95% CI 0–1.4%). Paid subscriptions bought no measurable safety gain.
The system being beaten here is not some app: it is the Australian government's own front door, treated as public health infrastructure and backed by a nurse phone line. The direction of the LLM errors matters too — when they undertriaged, they almost always pushed a self-care case toward a GP rather than keeping an emergency case at home, which is the expensive mistake rather than the dangerous one. The authors say plainly that this should feed regulatory discussion of whether LLMs are integrated with established symptom-checking services.
45 vignettes and 270 evaluations; the confidence interval on Healthdirect's accuracy spans 28 points, so either end is live. The Semigran set has been public since 2015 and is almost certainly in the training data — the largest caveat on the LLM side. Vignettes are not patients: no vague presenting complaint, no mid-consultation revision, no one who cannot describe their own symptoms. One country, one government product, one evaluation, and a preprint that has not been peer reviewed.
LUNA25: on 468 indeterminate lung nodules read by 75 radiologists, the best AI wins 0.78 to 0.69
Radiology: Artificial Intelligence vol 8 no 5 (September 2026) carries the results of the LUNA25 challenge run by Peeters, Obreja, Jacobs and colleagues at Radboudumc's Diagnostic Image Analysis Group (DOI 10.1148/ryai.260179). Training used 4,069 NLST CT exams; the hidden test set was 468 indeterminate nodules (156 malignant, 312 benign) drawn from the DLCST, NELSON and MILD trials, with a separate reader study in which 75 radiologists assessed 300 nodules. Per the authors' public abstract, the top system reached AUC 0.78 (95% CI 0.73–0.84, p<0.001) against a radiologist average of 0.69, finding 12% more malignancies at matched specificity and producing 20% fewer false positives at matched sensitivity. An accompanying commentary by Farina and Szarf ran on August 5.
Most "AI beats doctors" headlines come from retrospective sets and permissive comparisons. LUNA25 is not that: hidden test set, Docker submissions, the same nodules, the same humans, a pre-specified reader study. And it picked the expensive segment — indeterminate nodules. Neither obvious cancers nor obvious benign findings are the problem; sending 312 benign nodules into follow-up, biopsy or resection is where every lung screening programme actually bleeds. Twenty percent fewer false positives at matched sensitivity, at programme scale, is thousands of avoided workups and operations.
The winning algorithm is not a marketed product and holds no clearance anywhere. The readers worked in an artificial setting — no history, no priors, no clinical context — which handicaps the human side, since a radiologist in a real reading room has all three. Training came entirely from NLST (a US heavy-smoker cohort) and testing entirely from European trials, so representation of never-smokers and Asian populations is unknown. The Radiology: AI full text and commentary sit behind a paywall; the figures here come from the authors' public abstract and the challenge site.
Put the AI on the queue instead of the read: 43.4 minutes off the 90th-percentile wait for ED CT
The same issue carries a simulation study from Stanford emergency medicine and radiology with Harvard's biomedical informatics department — Silva, Tang, Chiacchia, Zhang, Langlotz and Kim (published July 22; DOI 10.1148/ryai.260110). They trained a gradient boosting model on 313,966 ED visits (August 2020–August 2024) to predict, at order time, whether a CT would prove clinically actionable, then ran discrete-event simulations comparing first-in-first-out against AI prioritisation across 20,795 studies from March–August 2024. For the 7,637 actionable studies (36.73%), median wait fell 10.75 minutes (95% CI −12.40, −9.10) and the 90th-percentile wait fell 43.36 minutes (−50.58, −36.80); completion within 60 minutes rose from 48.33% to 57.30%. The model captured 80–87% of the theoretical maximum benefit.
This is the most governance-friendly class of clinical AI on offer, because it never touches interpretation. The model does not say what is on the scan, only who should be scanned first; interpretive responsibility, liability and device classification all stay where they were, and no new hardware is needed. Set against the other six items here it reads almost as a counter-thesis: while everyone argues about what standard should validate diagnostic accuracy, this shows that pointing the prediction at scheduling routes the benefit around the entire validation deadlock.
This is a discrete-event simulation, not a prospective implementation; real ED scheduling is bound by staffing, scanner availability and patient flow that a simulation cannot capture. Wait time is also not a clinical outcome — nobody measured mortality, length of stay or changed management. And the benefit is not free: for non-actionable studies the median wait also fell 5.80 minutes, but the 90th percentile rose 14.46 minutes, meaning some patients were pushed further back. Single academic centre; generalisability untested.
252 marketed radiology AI products, 545 validation studies, and only 14% can say how they perform across subgroups
Shannon Walston, Hirotaka Takita, Yasuhito Mitsuyama, Junya Sato and corresponding author Daiju Ueda of the Department of Artificial Intelligence at Osaka Metropolitan University Graduate School of Medicine published a scoping review in European Radiology (DOI 10.1007/s00330-026-12652-y) covering 252 commercially available AI products from 142 manufacturers across 545 validation studies; it entered the news cycle this week via News-Medical on September 2 and Medical Xpress. Only 14% (77/545) reported performance stratified by sex, age or race/ethnicity. Eighty-one percent recorded patient sex but just 16% gave sex-stratified performance; 14% split patients into age bands, of which only 51% (40/79) included counts per band; race/ethnicity was lowest at 8% (45 studies), of which merely 18% (8/45) reported both demographics and performance. Skeletal (23% of 88) and lung (22% of 139) imaging most often included stratified data, and roughly 67% of tuberculosis detection datasets may lack the power for sex-based subgroup meta-analysis. Temporal analysis showed no improvement over time, and company sponsorship made no difference.
The first five items today answer the question of who the model beat. This one answers a different question: whether we are equipped to know who it fails. For 252 products already on sale, already cleared, already reading real patients' images, 86% of the validation literature cannot say. The finding that this has not improved over time is the ugliest part — it means several years of fairness guidance, reporting checklists and consensus statements have not changed what manufacturers actually put in the paper.
Radiology only, marketed products only — nothing on pathology, ECG or clinical decision support. The authors note that data extraction was done by a single reviewer, admitting possible selection bias; the review also does not compare performance between products or measure clinical adoption. The paper went online on June 12, 2026, so what happened this week was coverage rather than publication — it is here because it is the shared premise underneath the other six items.
One mammogram, a second organ system read out of it: AUROC 0.86 for stroke, 0.79 for hypertension, 0.78 for ischaemic heart disease
An August 27 press release from the European Society of Cardiology previewed a study presented at ESC Congress 2026 (Munich, August 28–31): Viana Copeland's group at Sheba Medical Center and Tel Aviv University trained a deep learning model on 29,921 women and 97,364 mammography examinations (median age 54) to identify hypertension (16% prevalence, AUROC 0.79), ischaemic heart disease (2.5%, 0.78) and stroke (2.5%, 0.86) from the mammogram alone. The ESC separately reported this as its largest congress ever, with nearly 34,000 professionals from 170 countries and more than 180 studies published simultaneously across journals (ESC).
Opportunistic screening is the cheapest form of clinical AI going: the image is already taken, the dose already delivered, the appointment already made — the model only reads out signal nobody was looking at. Copeland's own framing is that "because mammography is already widely used, analysing the same images for cardiovascular information could potentially offer a scalable approach without requiring an additional imaging examination." The link between breast arterial calcification and cardiovascular risk has years of literature behind it, but this measures three defined disease endpoints in nearly 30,000 women.
A congress abstract, with no peer-reviewed full text yet; every figure here comes from the ESC press release, and we did not obtain the original abstract tables. Single-centre retrospective data. And AUROC here measures identification of people who already carry the diagnosis, not prospective prediction of future events — clinically those are very different claims. At 2.5% prevalence, the positive predictive value behind an AUROC of 0.86 could be startlingly low, and the release does not report it.
02 — Product Analysis
Claude Opus 5 / Gemini 3.1 Pro
General-purpose frontier models as clinical knowledge engines · Anthropic (US) / Google DeepMind (US, UK)
Function and position. Neither is a medical device and neither holds a clinical indication anywhere, yet they scored 92.7% and 91.4% on specialist haematology board questions. They are sold to developers and enterprises; medical use largely arrives one of two ways — wrapped inside someone else's product (decision support, literature search, scribing), or as a tab a clinician opens privately. Neither route sits inside a device regulatory framework (medRxiv, 2026-09-02).
- Strength : consistently high scores across nine disease areas and six skill domains — a breadth no specialty AI product has — and, in a separate self-triage test, zero critical misses across 270 evaluations (medRxiv, 2026-09-02).
- Concern : multimodal scores drop 13–16 points, and half of haematology's diagnoses live in images. Worse, the top models fail on the same questions, so the commonest safety net — asking a second model — does not hold. The authors also flag that these public teaching banks are almost certainly inside the training data, and nobody can currently separate how much of the text score is reasoning and how much is recall.
- Missing : no public document reports how these models perform across patient sex, age or ethnicity — precisely the charge the European Radiology review levels at the commercial AI industry. General-purpose models are not even eligible for the charge, because they fall outside the 545 validation studies it counted.
Healthdirect Symptom Checker
Government-run online self-triage front door · Healthdirect Australia (federal and state funded)
Function and position. Healthdirect is Australia's publicly funded national health information and triage front door; its symptom checker walks users through rule-based questions toward self-care, a GP or an emergency department, with a 24-hour nurse line behind it. In this week's 45-vignette evaluation it triaged at 48.9% (95% CI 35.0–63.0), with Cohen's kappa 0.233, emergency sensitivity 46.7% and two critical misses (medRxiv, 2026-09-02).
- Strength : auditable, accountable, predictable. Every path through a rule-based system can be inspected and amended by a public health authority, and when it errs the failing rule can be traced. LLMs offer none of that today, and it is exactly why a government adopted this design.
- Concern : a kappa of 0.233 means agreement with correct triage is barely better than chance, and an emergency sensitivity of 46.7% means it misses more than half the cases that belong in an emergency department. For a front door carrying a government's name and read by the public as authoritative, that is a graver finding than anything charged against the LLMs in the same paper.
- Discount this : 45 vignettes cannot carry a verdict that the product is unsafe, and the confidence interval spans 28 points; the evaluators also used a seven-rule interaction protocol that may not match how the public actually uses it. The preprint has neither been peer reviewed nor answered by Healthdirect.
03 — Companies & Competition
| Company | Recent state & numbers | Position & moat |
|---|---|---|
| Qure.ai Chest X-ray and CT AI, India / US |
On February 26 it announced six further FDA 510(k) indications for qXR-Detect (lung, pleura, mediastinum/hila and heart, bone, hardware, other), taking it to 26 FDA clearances across nine products plus 65+ CE certifications, and making it the first chest X-ray CADe device cleared with a Predetermined Change Control Plan (Qure.ai, 2026-02-26). | The moat is regulatory breadth plus the update freedom a PCCP buys — rivals resubmit on every algorithm change; it does not. The thin spot is today's sixth item: tuberculosis detection is its core market, and that review found roughly 67% of TB datasets underpowered for sex-based subgroup analysis. |
| OpenEvidence Clinical decision support and medical search, US |
On September 3 it launched its own family of clinical models and deepened its oncology push (STAT, Fierce Healthcare); it had already signed a multi-year licence with JAMA to feed journal content into its AI medical search engine (Fierce Healthcare). | Its bet is content licensing plus clinician distribution rather than model capability: the haematology scores show knowledge itself is no longer scarce — citable provenance and an interface clinicians actually open are. The risk is that once general-purpose models secure equivalent licences, the moat reduces to habit. |
| Nihon Kohden Digital Health Solutions Continuous monitoring and deterioration prediction (CoMET), Japan / US |
Its cluster-randomised trial with the University of Virginia (an 85-bed cardiology medical-surgical ward, eleven clusters, 10,422 inpatient visits, January 2021–October 2022) published on February 5: the primary outcome, hours free of clinical deterioration within 21 days, showed no difference between arms, and the authors concluded that predictive analytics monitoring did not improve patient outcomes (Scientific Reports). | This row exists as a reminder: the only product in the table that ran a randomised trial got a null result. The moat is a data stream welded to monitoring hardware; the weakness is that it proved an accurate prediction is not the same as a patient who does better. Every AUC in today's report should be read against this line. |
| Anthropic / Google DeepMind / OpenAI General-purpose frontier models, no clinical indication |
On the 1,477-question haematology benchmark, Claude Opus 5 scored 92.7% on text, Gemini-3.1 Pro 91.4% and GPT-5.6 Sol 89.9%, while every model landed between 74.8% and 78.7% on multimodal items (medRxiv, 2026-09-02). A separate study the same week found paid tiers no safer than free ones on triage (medRxiv, 2026-09-02). | They have no moat because they never entered the market — other people's products carried them in. The real contest is not general versus specialist, but that a specialist vendor's entire validation and regulatory cost base is being hollowed out from the flank by a higher-scoring substitute bound by none of the same rules. Open-weight models improved most this generation, so that flank is widening. |
| Healthdirect Australia Government-run triage front door |
In the 45-vignette evaluation it triaged at 48.9% with kappa 0.233, emergency sensitivity 46.7% and two critical misses, trailing all six LLM configurations (medRxiv, 2026-09-02). | The moat is public trust and free universal access, not performance. Public trust is also the asset least able to survive this comparison: once people learn the free chat window triages better than the government's, the front door relocates itself — and there is no nurse line waiting on the other side. |
Today's competitive structure is an inverted pyramid. The vendor with the most clearances and the largest validation spend (Qure.ai) defends the narrowest indications; the only vendor that ran a randomised trial (Nihon Kohden) got a null result; and the layer scoring highest across the widest ground — general-purpose frontier models — is neither a device nor counted in any validation tally. This week, regulatory weight and benchmark performance line up in opposite order on the same table.
04 — Taiwan Angle
(1) Taiwan was the first country to extend LDCT screening to high-risk never-smokers — and its problem is precisely the indeterminate nodule. In the first year of the national programme, 49,508 people were screened (56.3% men, 43.7% women), 531 lung cancers were confirmed at a 1.1% detection rate, and 85.1% of those were stage 0 or 1. Among the eligible, 57.8% qualified on family history and 38.3% as heavy smokers; 74.6% of confirmed cases had a family history. Yet only about one in ten of an estimated 500,000 eligible people came forward (CommonHealth).
(2) Taiwan already has its own LDCT-AI, and it has been live for a while. The nodule-reading aid the Health Promotion Administration commissioned from National Taiwan University and engineering partners in 2021 received TFDA Class II device approval in February 2024 and is deployed at seven institutions: NTU Hospital, the NTU Cancer Center, National Cheng Kung University Hospital, Changhua Christian Hospital and the MOHW hospitals in Taipei, Nantou and Taitung. Official figures put detection accuracy at roughly 90% for nodules under 4 mm and over 6 mm with about two false positives, a 29% gain in report turnaround and a 31% gain in consistency of nodule characterisation (HealthNews).
(3) More consequentially, the National Health Insurance Administration is already moving AI toward reimbursement review. A March 2026 report describes an NHIA generative-AI recognition system being built as a pre-payment review platform for lung nodule surgery and as a self-management tool for hospitals. Division head Huang Yu-wen said the first-phase target is 95% agreement between the AI's assessment and the pathology report, that unnecessary operations on nodules judged benign would be denied, and that a central directive system would flag hospitals and physicians whose resection specimens come back malignant unusually rarely. The same report states that over 5.3 million LDCT examinations were performed in 2024 and nearly 30,000 people underwent nodule or tumour resection, with post-operative malignancy rates above 70% at NTU and three peer hospitals but below 50% — in places below 40% — elsewhere; clinical consensus is that nodules under 0.6 cm do not need surgery (UDN Health).
(4) And today's sixth item lands exactly on that intersection. If Taiwan is to deny payment on an AI's say-so, it must be able to answer one question: does this model perform the same in women, in never-smokers, in the over-60s, on scanners from different manufacturers? Only 14% of 545 global validation studies for commercial radiology AI can answer it, and only 8% even mention race or ethnicity (European Radiology). Taiwan's screened population is 43.7% women and enters mostly on family history rather than smoking — the very group NLST represented least, and NLST is where LUNA25's training data came from. Carrying an AUC of 0.78 from a model trained on American heavy smokers and validated on European trials over to Taiwan's screening population does not lack an algorithm. It lacks the subgroup validation nobody has done.
05 — Further Reading
-
What LUNA25 Teaches Us about AI for Lung Cancer Screening — Radiology: Artificial Intelligence (2026-08-05)
Farina and Szarf's accompanying commentary on LUNA25. After the headline AUC gap, this is the only peer-written document setting out what the challenge does and does not establish.
-
Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis — npj Digital Medicine (2026-01-28)
Pooling ten trials, it finds the counterintuitive result that clinician-plus-AI does not reliably beat AI alone: composite diagnostic and management scores gain just 4.88 points (95% CI 0.65–9.12) while factual error rates in documentation stay at 26–36%. Read it before assuming AI works as a second opinion.
-
The current state of demographic subgroup reporting for commercially available AI for radiology: a scoping review — European Radiology (2026-06-12)
The primary paper behind today's sixth item. If you need to judge whether any marketed imaging AI transfers to your own patients, this gives you the method, not just the verdict.
-
A benchmark study of vision and pathology foundation models for computational pathology — Nature Communications (2026-07-24)
Gevaert's Stanford group compares 32 foundation models across 41 tasks, 53 datasets and over 17,500 whole-slide images. Its conclusion rhymes with today's haematology benchmark: pathology-specific models are not uniformly better than general vision models, and scale does not guarantee generalisation.
-
Lessons from deploying the ChatEHR system at Stanford Medicine — Nature Medicine (2026-08-10)
A first-hand account of what actually happens once an LLM is wired into the record, centred on why benchmark scores are not a deployment case — precisely the question today's second and third items leave open.
06 — References
- Seuren LM, Williams R, Cresswell K, et al. Beyond productivity: a NASSS-informed review of implementation risks and research priorities for ambient AI scribes in healthcare. BMJ Digital Health & AI 2(1):e000091, 2026-09-03. DOI 10.1136/bmjdh-2026-000091. bmjdigitalhealth.bmj.com
- AI scribes may fail to capture patients' experiences. University of Edinburgh News, 2026-09-04. ed.ac.uk
- Ambient AI scribes may risk missing vital patient information. News-Medical, 2026-09-03. news-medical.net
- Radoynova M, Benouis M, Schulze F, Schneider MMK, Winter S, Bornhäuser M, Middeke JM, Eckardt JN. Benchmarking ten frontier large language models on 1,477 board-style multiple choice questions in hematology. medRxiv 2026.09.01.26361881, 2026-09-02. medrxiv.org
- Chowdhury AR, Chowdhury B. Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation. medRxiv 2026.09.01.26361908, 2026-09-02. medrxiv.org
- Peeters D, Obreja B, Antonissen N, Saghir Z, Pastorino U, De Bock G, Vliegenthart R, Prokop M, Jacobs C. Benchmarking of AI and Radiologists for Indeterminate Lung Nodule Malignancy Risk Estimation on Screening CT: The LUNA25 Challenge. Radiology: Artificial Intelligence 8(5), 2026-09. DOI 10.1148/ryai.260179. pubs.rsna.org
- LUNA25 Challenge, 作者公開版摘要與挑戰賽官網. Diagnostic Image Analysis Group, Radboudumc. diagnijmegen.nl · luna25.grand-challenge.org
- Farina EMJM, Szarf G. What LUNA25 Teaches Us about AI for Lung Cancer Screening. Radiology: Artificial Intelligence 8(5), 2026-08-05. DOI 10.1148/ryai.260657. pubs.rsna.org
- Silva E, Tang K, Chiacchia S, Zhang X, Langlotz C, Kim D. Simulation of AI-driven CT Queue Prioritization in the Emergency Department. Radiology: Artificial Intelligence 8(5), 2026-07-22. DOI 10.1148/ryai.260110. pubs.rsna.org
- Walston SL, Takita H, Mitsuyama Y, Sato J, Ueda D. The current state of demographic subgroup reporting for commercially available AI for radiology: a scoping review. European Radiology, 2026-06-12. DOI 10.1007/s00330-026-12652-y. link.springer.com
- Demographic reporting for medical AI remains inadequate, study shows. News-Medical, 2026-09-02. news-medical.net · Medical Xpress
- AI could help detect common cardiovascular diseases from mammograms. European Society of Cardiology press release, 2026-08-27. escardio.org
- Record-breaking ESC Congress 2026 puts new life-saving cardiovascular science in the spotlight. European Society of Cardiology, 2026. escardio.org
- Qure.ai nets six new indications cleared by the FDA. Qure.ai newsroom, 2026-02-26. qure.ai
- Keim-Malpass J, Ratcliffe SJ, Clark MT, et al. A randomized controlled trial of artificial intelligence-based analytics for clinical deterioration. Scientific Reports, 2026-02-05. DOI 10.1038/s41598-026-39051-z. nature.com
- OpenEvidence launches new family of AI models for clinicians. STAT News, 2026-09-03. statnews.com · Fierce Healthcare
- JAMA signs multi-year deal with OpenEvidence to inform AI-powered medical search. Fierce Healthcare. fiercehealthcare.com
- 晚期存活率剩1成!公費肺癌篩檢首年僅1成篩檢率. 康健雜誌 CommonHealth, 2023-07-03(2026-01-08 更新). commonhealth.com.tw
- 國健署LDCT肺癌篩檢AI輔助程式獲食藥署許可. 健康醫療網 HealthNews, 2024-09-07. healthnews.com.tw
- LDCT揪出一堆肺結節!健保署推AI把關,避免「不該切卻被切」. 聯合報元氣網 UDN Health, 2026-03-17. health.udn.com
- Wang G, Zhang K, Jiang J, Yang X, et al. Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis. npj Digital Medicine, 2026-01-28. DOI 10.1038/s41746-026-02382-2. nature.com
- Bareja R, Carrillo-Perez F, Zheng Y, Pizurica M, et al. A benchmark study of vision and pathology foundation models for computational pathology. Nature Communications, 2026-07-24. DOI 10.1038/s41467-026-76004-6. nature.com
- Lessons from deploying the ChatEHR system at Stanford Medicine. Nature Medicine, 2026-08-10. nature.com
- Lång K, et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study. The Lancet, 2026-01-29. DOI 10.1016/S0140-6736(25)02464-X. eurekalert.org