Right Answer, Wrong Route: Three Medical-AI Scorecards This Week All Broke on the Same Question — Can You Audit the Path?
A technology column usually reports how many points the score went up. Not this week. On August 7, CliniCARE-Bench put 16 agentic systems through 750 audit scenarios derived from real patient data. Four-way accuracy landed at 65.3–76.1% — respectable, until the authors computed a second number: defect-free accuracy, which credits a verdict only when it is both correct **and** free of prohibited shortcuts. That figure fell a further 4.8 to 14.8 points. The same month, EHR-Complex measured a best-in-class 62.3% exact match over 52K tasks, yet Pass^k consistency dropped below 50% for nearly every model at k=4. And on August 17 Nature Biomedical Engineering published ConceptCLIP, which stops treating explanation as a post-hoc add-on and writes it into the pretraining objective — 23 million image–text–concept triplets bought in exchange for a model a clinician can check line by line. Three scorecards, one message: the frontier of medical AI has moved from *is it accurate* to *can you audit it*.
01 — Top Stories
ConceptCLIP: 23 million image–text–concept triplets turn interpretability from a post-hoc patch into a pretraining objective
On August 17, Hao Chen's group at the Hong Kong University of Science and Technology published ConceptCLIP in Nature Biomedical Engineering, releasing alongside it MedConcept-23M — 23 million biomedical image–text–concept triplets. Where ordinary contrastive vision–language pretraining aligns an image with a block of text, ConceptCLIP adds region–concept alignment, so the model can point to the part of the image it relied on and name the medical concept that region maps to. Evaluation spans 78 datasets across 10 imaging modalities — chest X-ray, CT, MRI, mammography, whole-slide histopathology, endoscopy, ultrasound, dermatology, retinal imaging and more — comprising 68 public datasets, 9 private ones, and a newly curated PMC-9K set of 9,222 image–text pairs. Weights are on Hugging Face, code on GitHub, and the dataset captions at Hugging Face Datasets.
For two years, interpretability in medical imaging foundation models has been almost entirely post-hoc: Grad-CAM heat maps, attention visualisations, or asking an LLM to narrate its reasoning. All share one flaw — the explanation and the decision come from different places. How the model computed its answer is one thing; the story generated afterwards is another. ConceptCLIP puts concept alignment inside the loss function, making the explanation part of the representation itself. The paper's accompanying clinician user study, run across three imaging modalities, found that concept-based explanations helped clinicians verify predictions and spot potential errors — a more interesting result than any benchmark number, because it answers whether the explanation is *useful*, not merely whether one exists.
Nine of the 78 datasets are private, so that portion of the results cannot be fully reproduced externally. The user study covers three modalities, not all ten. Parameter count and training compute are not disclosed in the accessible material, and what MedConcept-23M releases on Hugging Face is the captions rather than every source image — rebuilding the full corpus requires sourcing the images yourself.
CliniCARE-Bench: 16 agentic systems get roughly seven in ten right — until shortcut answers are subtracted, costing up to 14.8 points
A multi-institution team including Stanford's Nigam H. Shah released CliniCARE-Bench (arXiv:2608.07796) on August 7. It takes 25 clinician-validated retrospective audit scenarios and instantiates them as 750 patient-specific cases derived from real MIMIC-IV data, asking each system for one of four verdicts: Yes, No, Indeterminate (Lack of Data), or Indeterminate (Medically Ambiguous). Across 16 agentic systems, four-way accuracy ranged from 65.3% to 76.1%. The paper then defines defect-free accuracy — credit only when the verdict is correct *and* the reasoning avoided explicitly prohibited shortcuts — and that figure comes in 4.8 to 14.8 points lower.
Two design choices deserve to be copied. First, splitting "not enough data" from "medically ambiguous" into separate verdicts: most medical QA benchmarks offer only right or wrong, forcing a guess when information is missing and thereby rewarding confident error. Second, defect-free accuracy moves evaluation from *was the output correct* to *was the process acceptable*. That is precisely the difference between an audit and a quiz — an audit demands a reviewable chain of reasoning, not a conclusion that happens to be right. Read it against the Microsoft Research and Scripps finding published earlier this year in Nature Medicine, where frontier models held their accuracy even after critical visual information was removed — meaning they were not looking at the image at all — and the two results are the same lesion in different contrast.
This is an arXiv preprint and has not been peer reviewed. The abstract does not name the 16 systems or break out individual scores, so the public summary gives no way to tell which vendor lost the most. All 25 scenarios derive from MIMIC-IV — critical-care data from a single US academic centre — and transferability to Taiwanese or European health systems is untested.
EHR-Complex: 52K tasks, 3,800 failure traces, a 62.3% ceiling — and consistency below 50% when you run it four times
Yitong Qiao and colleagues built EHR-Complex (arXiv:2606.23301) on top of MIMIC-IV — 365,000 patients, 31 tables, over 500 million records — generating roughly 52,000 tasks across six clinical intents, supporting both patient-level and population-level queries. Agents must actually execute SQL or Python in a sandbox to answer. Top exact-match performance is 62.3%. More telling: Pass^k consistency drops below 50% for nearly all evaluated models at k=4. Analysis of more than 3,800 failed attempts sorts the errors into three buckets: SQL logic mistakes, medical-code lookup failures, and semantic misreadings.
Pass^k is the number worth citing. It asks not "does one run get it right" but "does the same question come out right every time across k runs". A 62.3% single-shot score looks usable in a demo video, but a system that answers inconsistently on two of four attempts cannot compute a hospital quality metric or audit a reimbursement claim — those jobs need reproducibility, not an expected value. It also explains why medical-code lookup earns its own failure category: the hierarchy relations in ICD-10 and SNOMED are not something a language model can infer from general knowledge, they must be looked up, and models are unusually good at papering over a missing lookup with fluent prose.
Also an unreviewed arXiv preprint, submitted June 22. Tasks are programmatically generated, so their distribution may not match what clinical staff actually ask. The author list is weighted toward Ant Group and Zhejiang University; whether the evaluated model set includes their own systems requires reading the full paper.
The same explanation helps the expert and fools the novice: Nature Medicine shows XAI has no universal setting
Led by Orson Xu of Columbia's Department of Biomedical Informatics, with MIT's Marzyeh Ghassemi and Stanford's Roxana Daneshjou among the co-authors, this Nature Medicine paper of August 4 tested explainable AI in dermatological diagnosis. Non-experts and primary care physicians diagnosed skin conditions from images, with and without various XAI aids — confidence scores, similar-image retrieval, heat maps, and LLM-written explanations. The two groups diverged sharply. Lay accuracy rose, but by deference: they trusted LLM explanations regardless of correctness, and found vaguer explanations more convincing. Clinicians did best with the bare prediction and no explanation at all — they caught the model's errors and were not led astray by incorrect assistance. MIT News quotes the researchers' conclusion: "the same explanation can help an expert and mislead a beginner."
This is a counterpunch to the two benchmarks above. CliniCARE-Bench and ConceptCLIP both assume better explanations mean safer systems; this paper shows the effect depends entirely on the receiver's ability to push back. For someone equipped to disagree, an explanation is an auditing tool. For someone who is not, it is a persuasion tool. The design implication is blunt: the clinician-facing and patient-facing renderings of the same model output have to be two different interfaces, not the same one in a different font. And with OpenAI having wired ChatGPT Health straight into patient portals, where the overwhelming majority of users are non-experts, "vaguer explanations were more convincing" is the sentence in this study that should keep product managers awake.
The engineering half of Epic UGM 2026: a generative model on 320 million patients, and an agent factory handed to people who cannot code
Epic held its annual Users Group Meeting in Verona, Wisconsin on August 17–20 (20,000-plus attendees across roughly 500 sessions; Fierce Healthcare counts 8,400 in person and around 70,000 including virtual). Three engineering outputs: Curiosity, a generative model trained on Cosmos data to forecast outcomes such as readmission and stroke risk, available to data-contributing Cosmos members from March 2027, with Cosmos now covering 320 million patients and 23 billion encounters; Agent Factory, a no-code platform for building and orchestrating agents with 120-plus pre-built ones included, going broadly available in 2027, where early adopter ECU Health built a transfer-request summarisation agent estimated to save 20 hours a week; and Ergo / Ergo Visit, an AI-native clinical interface shipping November 2026, whose visit-prep function draws on 320 million de-identified records.
Set Curiosity beside the three items above and an awkward mismatch appears. Academia is pushing evaluation toward auditable process and run-to-run consistency; the industry's largest single bet is a generative risk-prediction model trained on a corpus too large to trace case by case. That is not to say Epic is wrong — longitudinal data on 320 million people is an asset perhaps three organisations on earth possess — but when Curiosity begins emitting readmission risk scores in March 2027, CliniCARE-Bench's defect-free framing lands squarely on it. Agent Factory's exposure is more concrete: 120 pre-built agents, no code required, each hospital customising its own, distributes the agent safety boundary from one vendor's responsibility across thousands of hospital configuration screens — and EHR-Complex's Pass^k result already tells us the same agent run four times need not give the same answer.
ECU Health's "20 hours a week" and Ask Emmie's "1,000 call-centre hours saved" at Sutter Health are customer-reported and not independently audited. Neither Curiosity nor Agent Factory has shipped (March 2027 and 2027 respectively), so every description comes from vendor presentations with no independently verified performance figure attached. The two attendance counts conflict (20,000 versus 8,400 in person); both are listed here rather than reconciled.
Suki goes the other way: in the age of ambient scribing, rebuilding plain dictation as an API
AI scribe company Suki launched Suki Dictation on August 25 — natively built inside Epic and Meditech, deployable standalone or alongside Suki for Clinicians, and offered via API and SDKs for third-party health tech products to embed. Unlike ambient documentation, which passively records the encounter, this lets clinicians actively edit by typing or voice command, placing text exactly where they want it in the chart. Pilot user Edward Mintz, a pulmonary and critical care physician at Witham Medical Group, reports saving at least one to two hours a day. Chief product officer Abhi Pathak framed it this way: "We are not trying to win by locking people into a bigger bundle. We are trying to give health system clinicians the freedom to choose exactly what they need." The company also previewed next-generation streaming speech recognition, conversational voice editing and AI-assisted procedure notes.
This is a technical fork worth noticing. The dominant narrative of the last three years has been that ambient AI would eat dictation: the clinician does nothing, the AI listens and drafts. Suki is betting the opposite way — that when a generated draft still needs heavy editing, the bottleneck is not generation but precise correction. Turning dictation into an API concedes that voice is an atomic operation inside the clinical workflow rather than a complete product. It also explains the refusal to bundle: with Epic shipping Ergo and Agent Factory, any third party competing head-on for the whole workflow is fighting uphill, and becoming the component everyone else has to license is a defensible place to stand.
The "one to two hours a day" figure comes from a single pilot physician — sample size one, with no control arm or time-and-motion study. The coverage discloses no pricing, no word error rate, no supported languages, and no count of contracted organisations.
A JAMA viewpoint argues that once AI clearly beats physicians, mandating a human in the loop locks in worse care
A JAMA viewpoint (DOI 10.1001/jama.2026.15380) with Ezekiel J. Emanuel of the University of Pennsylvania's Department of Medical Ethics and Health Policy as corresponding author appeared on August 17 under the title "Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care?", with an EurekAlert release the same day. The Decoder's August 18 summary lists the evidence cited: Google's AMIE outperformed primary care doctors in simulated consultations; ChatGPT o3 reached the correct diagnosis first in 60% of complex cases against 15.9% for internists; GPT-4 alone scored 92% on diagnostic reasoning while physicians with AI access scored 76%. The piece argues that rules requiring physician sign-off risk institutionalising inferior care by 2030. Co-authors include Curai Health CEO Neal Khosla, whose father invests in both OpenAI and Curai.
Read alongside today's dermatology XAI paper, an uncomfortable question surfaces. Emanuel and colleagues' "GPT-4 alone 92%, doctors with AI 76%" and Nature Medicine's finding that clinicians perform best with a bare prediction and worse once explanations are added may be two readings of the same phenomenon: the loss in human–AI collaboration need not come from humans dragging the model down, it may come from interfaces putting the human in the wrong place. The two attributions point to opposite policies — remove the human, or rebuild the interface. Given that nearly all the evidence comes from simulation rather than delivered care, the first is a far harder bet to unwind than the second.
This is a viewpoint, not primary research. The authors themselves concede that nearly all the evidence comes from simulations rather than delivered patient care, and that invasive procedures still require physicians. Two authors hold direct financial interests in the spread of autonomous AI — a conflict to discount for while reading. The JAMA text sits behind a paywall; the figures in this item are taken from the EurekAlert release and The Decoder's secondary summary, not from the full article.
Six Nordic countries turn national longitudinal health data into a federated AI platform — and publish the blueprint in Nature Medicine
The Nordic AI-Health Initiative paper, led by Ole A. Andreassen with 22 authors, appeared in Nature Medicine on August 14 under the subtitle "technical foundations, data assets, and a roadmap for deployment". Authors span Norway, Sweden, Denmark, Finland, Iceland and Estonia. The goal is a federated Nordic health ecosystem: rather than centralising data, models travel to where the data sits, under regulation-compliant access controls, to support development of generalisable medical AI.
What matters here for a Taiwanese reader is not novelty — federated learning is not a new idea — but sequencing. The Nordics fixed cross-border data governance and standards first, then talked about models; that is very nearly the path Taiwan's Ministry of Health has taken by building FHIR Box before deploying AI applications (see the Taiwan section). Epic's Cosmos, by contrast, concentrates 320 million people's data with one vendor and wins on scale. The difference between the two architectures will show up concretely in two places over the next few years: who can build models that generalise across systems, and who is answerable when a model errs — centralised gives you one accountable party, federated distributes responsibility across each participating country's governance framework.
The paper is a blueprint and roadmap; the publicly accessible portion does not state the exact record counts of each country's data assets, nor list dated milestones. No unverified scale figures are quoted here.
02 — Product Analysis
ConceptCLIP
Concept-aligned, interpretable biomedical vision–language model · HKUST (Hong Kong)
Function and position. ConceptCLIP is a general-purpose biomedical image encoder built so that a clinician can check it, not so that it tops a leaderboard. It performs joint image–text and region–concept alignment over MedConcept-23M (23 million image–text–concept triplets), evaluated across 78 datasets and 10 imaging modalities. The buyer is a hospital IT department or a team deploying imaging AI in a regulated setting — anywhere you must explain to an auditor what the model relied on.
- Strength : explanation and decision share a source. Concept alignment lives in the training objective rather than being narrated afterwards, and a clinician user study across three modalities found physicians could use it to catch the model's errors.
- Strength : weights and code are public (GitHub, Hugging Face), so it can be deployed inside a hospital's own environment with no data leaving the building.
- Concern : nine of the 78 evaluation datasets are private, so the results cannot be fully reproduced externally, and with no parameter count or training compute disclosed, an adopter cannot size the hardware. Today's Nature Medicine dermatology study is also a caution: useful to a clinician does not imply useful to a patient.
MedGemma 1.5
Google's open-weight medical multimodal model · Google (US)
Function and position. MedGemma 1.5 sits on top of Gemma 3, updated January 13, 2026, currently shipping as a 4B multimodal variant covering 3D imaging (CT/MRI volume reading), whole-slide pathology, longitudinal image comparison, anatomical bounding boxes, and clinical document and EHR understanding. Figures from the official model card: 61.1% macro accuracy on CT (up from 58.2%), 64.7% on MRI (up from 51.3%), 89.5% macro F1 on MIMIC chest X-ray, and 89.6% on EHRQA. The buyer is a developer — this is the entry point to the Health AI Developer Foundations ecosystem.
- Strength : 3D CT/MRI and whole-slide pathology at 4B parameters puts the deployment bar well below comparable models, and the jump on MRI from 51.3% to 64.7% is the most substantial gain in this generation.
- Concern : interpretability remains post-hoc, and the model card carries no user study on whether a clinician can actually check its work. Licensing is governed by the Health AI Developer Foundations terms rather than a standard open-source licence — read it clause by clause before commercial deployment.
- Concern : accuracy in the 61–65% band is a long way from unsupervised operation in a radiology department, and the Nature Medicine robustness study argues that what this class of VQA benchmark actually measures is itself contested.
03 — Companies & Competition
| Company / group | Recent state & numbers | Position & moat |
|---|---|---|
| Epic Systems Largest US EHR vendor |
UGM 2026 unveiled Curiosity (March 2027), Agent Factory (2027, 120+ pre-built agents) and Ergo (November 2026). Cosmos spans 320 million patients and 23 billion encounters; 43.7% acute-care EHR share, 3,700-plus hospitals, records for 325 million people (Fierce Healthcare). First German customer Charité signed at roughly €200 million (HealthTech HotSpot). | The moat is data scale plus workflow lock-in; nobody else can assemble comparable longitudinal data. The thin spot is auditability: Cosmos is too large to trace training provenance case by case, exactly as academic evaluation standards move the other way. |
| Google Open-weight medical model supplier |
MedGemma 1.5 (updated January 13, 2026), 4B multimodal: 61.1% on CT, 64.7% on MRI, 89.5% macro F1 on MIMIC CXR, 89.6% on EHRQA (official model card). AMIE beat primary care physicians in simulated consultations and is cited in the JAMA viewpoint as evidence for autonomous AI (The Decoder). | The play is free models in exchange for ecosystem position: get the world's health AI startups building on HAI-DEF, then collect on cloud and tooling. The weakness is a non-standard licence and no evidence base for clinician-checkable output. |
| HKUST / ConceptCLIP Academic group, open weights |
MedConcept-23M (23 million triplets), 78 datasets, 10 modalities, PMC-9K (9,222 pairs), with weights and code released (Nature BME, 2026-08-17). | It does not compete on parameters but on checkability — a differentiator with real value under the EU AI Act and FDA's approach to generative devices. But an academic group has no operations, versioning or compliance support; real deployment will need a vendor to wrap it. |
| Suki AI scribe and voice components |
Launched Suki Dictation on August 25, native inside Epic and Meditech and exposed through API and SDKs for third-party integration; a pilot physician reports saving one to two hours a day (Fierce Healthcare, 2026-08-25). | Deliberately declines to fight Epic for the workflow, positioning instead as the voice component everyone licenses. The moat is thin — speech recognition is commoditised — but the distribution position, native inside two major EHRs, is hard to replicate quickly. |
| 長佳智能(6841)/遠傳(4904) Taiwan medical AI and systems integration |
Long Chia Intelligence holds around 57 domestic and international medical device licences, with H1 2026 revenue of NT$167.1 million and gross margin rising from 57.8% to 73.1%; FarEasTone is the principal systems integration partner for the FHIR-based smart healthcare data platform (Yam News, 2026-08-24). | The moat comes from a policy timetable rather than technology: FHIR Box's three phases — medical centres in 2026, regional and district hospitals in 2027, clinics and health stations in 2028 — determine who plugs into the pipe first. The risk is symmetric: if the timetable slips, valuations take the hit before deployments do. |
Close the section in one sentence: today's competitive structure is two axes — data scale and checkability — pulling apart. Epic and Google each hold a position on the scale end, one on longitudinal data covering 320 million people, the other trading free weights for ecosystem position. Academic outputs like ConceptCLIP bet on the other end, wagering that regulation and clinical adoption will eventually demand a stated basis for every call. Taiwan's position is unusual: it has neither Cosmos-scale data nor a frontier model team, but by fixing data standards before applications, its health ministry has landed on the checkability side of the split.
04 — Taiwan Angle
(1) What FHIR Box is for is precisely the gap today's three benchmarks expose. The Ministry of Health's FHIR Box is a universal operating layer for healthcare, letting hospitals exchange standardised records without altering their existing systems, so that AI models can work across institutions rather than being customised for each one. The rollout runs in three phases: record interoperability across medical centres by end-2026, regional and district hospitals by end-2027, clinics and health stations in 2028 (Yam News, 2026-08-24). The failure modes CliniCARE-Bench and EHR-Complex surface — medical-code lookup errors, unstable multi-table longitudinal reasoning — are at root problems of semantic inconsistency in the data, which is exactly the layer FHIR standardisation addresses. Taiwan's technical opening is not building a bigger model; it is becoming one of the few places able to offer a standardised, traceable, cross-institution evaluation environment.
(2) On August 21, the health ministry signed an MOU with Roche Diagnostics, targeting chronic kidney disease first. The work centres on AI early warning to slow deterioration, against a backdrop of more than 90,000 dialysis patients in Taiwan. What stands out is the sequencing ministry officials emphasised: establish data standards and governance first, deploy AI tools second. Chang Gung, Mackay and Chung Shan medical centres already share records in real time through FHIR Box to support clinical AI applications (CNA, 2026-08-21). That ordering matches the Nordic AI-Health Initiative almost exactly and runs opposite to Epic's centralised Cosmos — for a system with complete national insurance data but limited scale at any single institution, federated is the pragmatic choice.
(3) But today's XAI study is a warning for Taiwan's patient-facing plans. Nature Medicine concluded that non-experts are persuaded by vague explanations. If the chronic disease AI early-warning work eventually reaches patients — pushing kidney-function risk through a health passbook or an app — the interface has to assume a recipient who cannot argue back, which makes it a different product from the nephrologist's view of the same model. Separately, Long Chia Intelligence's H1 2026 revenue of NT$167.1 million at 73.1% gross margin (Yam News) reflects a policy-dividend window; any investment read should discount for schedule risk, since a slip in any one of FHIR Box's three phases hits policy-timetable-priced names first.
05 — Further Reading
-
Medical AI has a measurement problem — Nature (2026-07-28)
Harvard Medical School's Arjun K. Manrai distils the anxiety behind all three of today's benchmarks into one line: the technology is outrunning our ability to measure it. Read it first and the design motives behind defect-free-style metrics become much clearer.
-
Evaluating the robustness and readiness of large frontier models in health AI applications — Nature Medicine (2026-06-26)
A 32-author Microsoft Research and Scripps team found that removing critical visual information barely dents frontier model accuracy. This is the empirical starting point for every "right answer, wrong route" argument in today's issue.
-
The benefits of medical AI assistance vary based on user expertise — MIT News (2026-08-04)
The plain-language version of the Nature Medicine dermatology study, clearer than the abstract on experimental design and the four XAI conditions. Anyone building patient-facing products should read this one directly.
-
Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care? — JAMA (2026-08-17)
The most-cited piece of the year, and the one most in need of being read together with its conflict-of-interest statement. Agree or not, it puts "is a human in the loop a protection or a drag" formally on the table. Paywalled — start with the EurekAlert release.
-
An AI-Health infrastructure for the Nordic region — Nature Medicine (2026-08-14)
The most directly useful piece for Taiwan: cross-border health data governance written up as an engineering specification. To understand what FHIR Box will run into over the next three years, the fastest route is someone else's already-written roadmap.
06 — References
- An explainable biomedical foundation model via large-scale concept-enhanced vision–language pretraining. Nature Biomedical Engineering, 2026-08-17. nature.com
- ConceptCLIP — code repository. GitHub. github.com
- MedConcept-23M dataset captions. Hugging Face Datasets. huggingface.co
- CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR. arXiv:2608.07796, 2026-08-07. arxiv.org
- EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning. arXiv:2606.23301, 2026-06-22. arxiv.org
- Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people. Nature Medicine, 2026-08-04. nature.com
- The benefits of medical AI assistance vary based on user expertise. MIT News, 2026-08-04. news.mit.edu
- Epic expands AI ambitions with agent platform, Cosmos-powered predictions and deeper workflow automation. Fierce Healthcare, 2026-08. fiercehealthcare.com
- Epic UGM 2026 Recap: Cosmos Curiosity, EpicOps and an Expanding AI Footprint. HealthTech HotSpot, 2026-08. healthtechhotspot.com
- AI scribe Suki rolls out AI-driven clinical dictation tool. Fierce Healthcare, 2026-08-25. fiercehealthcare.com
- Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care? JAMA, 2026-08-17. DOI 10.1001/jama.2026.15380. jamanetwork.com
- Will autonomous AI exceed AI-aided physicians as the best medical care? EurekAlert!, 2026-08-17. eurekalert.org
- As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece says. The Decoder, 2026-08-18. the-decoder.com
- An AI-Health infrastructure for the Nordic region: technical foundations, data assets, and a roadmap for deployment. Nature Medicine, 2026-08-14. nature.com
- Evaluating the robustness and readiness of large frontier models in health AI applications. Nature Medicine, 2026-06-26. nature.com
- Medical AI has a measurement problem. Nature, 2026-07-28. nature.com
- MedGemma 1.5 model card. Google Health AI Developer Foundations, 2026-01-13. developers.google.com
- MedGemma: Our most capable open models for health AI development. Google Research Blog. research.google
- Introducing ChatGPT Health. OpenAI. openai.com
- 衛福部攜手羅氏推 AI 醫療 首波瞄準慢性腎臟病照護. 中央社 CNA, 2026-08-21. cna.com.tw
- 台灣醫療 AI 掀全國基建!FHIR Box 串聯跨院 長佳智能、遠傳受惠. 蕃新聞 Yam News, 2026-08-24. n.yam.com
- STAT Health Tech: FDA promises new AI guidance, and Epic UGM updates. STAT News, 2026-08-25. statnews.com
- This Week in European HealthTech, MedTech and Health AI: 21st August 2026. healthcare.digital, 2026-08-21. healthcare.digital