◆ AI & Medical AI Daily
–
Thursday · Technology Breakthroughs

Every technical breakthrough in medical AI this week ran into the same wall — the scorer is the vendor: OpenEvidence raised another $250M at a $15B valuation on 9/30, and the first-ever 100% on MedQA that its Darwin model posted was graded by OpenEvidence itself; a Nature Medicine head-to-head found three general-purpose frontier LLMs beating two leading specialized clinical AI tools on MedQA, HealthBench and real physician questions, with the specialized tools doing no better than a Google AI overview; and Microsoft's own HealthAgentBench put 54 agentic tasks to frontier models, where the best — Codex GPT-5.5 — cleared just 42%

Thursday is the technology slot, and today opens with an admission: between 9/28 and 10/1 no flagship medical model shipped. What actually landed were arXiv preprints and weights appearing on Hugging Face — of the five "new models" swept up by the 9/30 medical AI industry digest, the most downloaded, Fastino-Nemotron-3.5-Lightning-Healthcare, sits at 6,707 downloads and 22 likes; the other four together clear fewer than 2,000, and none carries an independent clinical-accuracy evaluation. So today's technical thread is not whose model is stronger. It is a harder question: as the scores keep climbing, who grades them independently? On 9/30 that question acquired a price tag — OpenEvidence raised another $250 million at a $15 billion valuation, and the four numbers behind its claim to the world's highest-scoring medical AI model family on every mainstream benchmark were measured, graded and published by OpenEvidence.

01 — Top Stories

Seven items in one line: the scores, the valuation, the real pass rate on agentic tasks — and the one kind of progress that needs no scoreboard argument
Funding / Benchmarks OpenEvidenceDarwin9/30

OpenEvidence raised $250M at a $15B valuation on 9/30 — and the 100% on MedQA that its Darwin model posted was a score it gave itself

What

On September 30, OpenEvidence — widely described as ChatGPT for doctors — closed a $250 million round at a $15 billion valuation from investors including health systems and Andreessen Horowitz, bringing its total to roughly $1 billion over twelve months. The technical case behind the round came from the OpenEvidence Model Family announced on September 3: Darwin became the first AI system in history to score 100% on MedQA, alongside 72.8% on MedXpertQA, 82.7% on HealthBench Professional and 87.2% on NOHARM. The company says Darwin beat the equivalent Claude and Gemini models, but it remains in research preview over dual-use risk in virology, immunology and genetics. The other three — Osler (5-second answers), Sackett (30 seconds, evidence-led) and Snow (5-minute literature review) — are free to US clinicians.

Why it matters

Because nobody independently verified those four numbers. A September 22 analysis by Grid Health put it plainly: "these are OpenEvidence's own tests, scored by OpenEvidence, published by OpenEvidence." The sharper cut is the naming — Sackett honours David Sackett, the father of evidence-based medicine, while supplying only self-generated evidence; the same piece notes that OpenEvidence had criticized NYU researchers for using OpenAI benchmarks, yet HealthBench Professional, the benchmark now anchoring its world-first claim, is an OpenAI benchmark. This is not a doubt about capability. It is a doubt about the gap between clinical accuracy and clinical utility: 100% on MedQA is not the same as doing the right thing in a consultation room.

Discount this

All four benchmark figures are vendor-reported and not independently audited. The company's claim that more American physicians use its platform than every competing AI tool combined comes with no citation. The round size and valuation are from secondary coverage at HISTalk, which itself flags that earlier reporting had the company seeking $200 million at $20 billion — numbers that do not match the close.

Paper / Head-to-head Nature MedicineNYUMedQA · HealthBench

The Nature Medicine verdict: general-purpose frontier models beat specialized clinical AI tools — and the specialized tools did no better than a Google AI overview

What

Krithik Vishwanath, Anton Alyakin and Eric Karl Oermann did something this industry rarely does: they sat three general-purpose frontier LLMs at the same table as two leading specialized clinical AI tools and compared them head to head across two public benchmarks — MedQA and HealthBench — plus questions actually submitted by physicians. All three general-purpose models won, and the two specialized clinical tools performed no better than Google search's AI overview. The research briefing ran in Nature Medicine vol. 32, no. 7, pp. 2364–2365 (DOI 10.1038/s41591-026-04457-9, 2026-06-17); the full paper, "General-purpose large language models outperform specialized clinical AI tools on medical benchmarks," is indexed at PubMed 42286322 / s41591-026-04431-5.

Why it matters

The risk the study names is exactly the risk in the 9/30 round: specialized clinical AI tools are entering medical practice with little independent testing. Stack the two together and the landscape looks like this — general-purpose models are now good enough at general medicine that building a medical-specific model is a proposition requiring proof, and proof requires someone outside the vendor to do the grading, and that someone does not currently exist. Arjun K. Manrai of Harvard Medical School's Department of Biomedical Informatics compressed it in a July 28 Nature commentary: evidence-based medicine rests on the assumption that the yardstick is reliable, yet medical AI capability is "increasingly difficult to assess and understand properly."

Discount this

This is not this week's news — the paper published in June 2026. It is pulled back in today because it is the only independent yardstick against which this week's funding round and self-graded scores can be read. Also, the Nature Medicine full-text page returned a 503 at the time of writing; this item is built from the research briefing and the PubMed record, not the full text. The briefing itself does not name which three general-purpose models or which two specialized tools were tested, and publishes no per-model scores.

Benchmark / AI agents Microsoft ResearchHealthAgentBenchCodex GPT-5.5

Microsoft put 54 real healthcare agent tasks to frontier models; the best, Codex GPT-5.5, cleared 42% — and medical imaging was where they struggled most

What

HealthAgentBench (arXiv 2606.31179, submitted 2026-06-30), led by Qianchu Liu and Sheng Zhang with Hoifung Poon as corresponding author, assembles 54 agentic healthcare tasks across 7 categories, each replicating a real clinical workflow: the agent must explore raw healthcare data itself, operate in a complex environment and execute a multi-step solution rather than answer a prompt. The result: Codex GPT-5.5 leads at roughly 42% success, and is also the most cost-effective; the Claude Code models struggled notably on medical imaging; and every frontier agent faltered on tasks combining a large search space with compositional reasoning. The benchmark is publicly released as microsoft/HealthAgentBench on GitHub.

Why it matters

Put 42% next to 100% on MedQA and you have the central gap in today's report. The ceiling on multiple choice has been broken through; the floor under workflow has not been laid. MedQA is one question, one answer, verifiable, no side effects. A HealthAgentBench task is "go into this pile of raw data, find what needs finding, then get the next five things right" — and the latter is what hospitals actually buy. Worth noting too: Microsoft's own benchmark marks down models from Microsoft's own partners. That is what independent evaluation looks like, whoever builds it. Agents showed promise on EHR-based research pipelines but failed collectively on imaging, which is also why imaging AI remains single-purpose models plus a radiologist's signature rather than an agent running the whole thing.

Discount this

Also not this week's news (submitted June 30). The 42% is approximate; the paper does not give a full leaderboard at abstract level, and per-model scores for the Claude Code family and the other frontier agents were not obtained for this item. Fifty-four tasks is a thin basis for a benchmark meant to represent an entire clinical workflow, and the authors themselves frame the low pass rates as the benchmark being hard rather than the models being bad — two readings the data cannot currently separate.

Open evaluation SophontMedmarks v1.030 benchmarks

Sophont's Medmarks v1.0 lays out 30 benchmarks across 61 models: medical fine-tuning genuinely helps, large open models burn 5x the tokens to keep up — and handing a model a calculator made it worse

What

Sophont expanded its medical LLM evaluation suite to 30 benchmarks (from 20) and 61 models across 71 configurations (from 46), releasing a technical report (arXiv 2605.01417), an interactive leaderboard and the Medmarks-T training environment — all open source. On the verifiable split, Medmarks-V, Gemini 3 Pro Preview leads; on open-ended clinical reasoning, Medmarks-OE, GPT-5.2 ranks highest. Three findings are worth copying down: (1) medical fine-tuning genuinely works — fine-tuned variants achieve near-Pareto improvements over their base models across families; (2) the open-weights advantage holds only on multiple choice — GLM 4.7 approaches frontier performance on MCQs, loses that edge on harder open-ended tasks, and needs "over 5x the number of tokens" for results comparable to GPT-5.2, a real problem for clinical deployment where speed matters; (3) tool integration backfires — given a calculator, many models ignored it, produced wrong answers despite having it, or suffered formatting failures once an external tool was in play.

Why it matters

Medmarks does the thing OpenEvidence did not: it hands the grading to someone other than the vendor. Its second finding lands on a cost most people skip — "open models caught up on the scoreboard" is not "open models are deployable." When a model needs 5x the tokens to draw level, in a hospital running tens of thousands of queries a day that 5x is inference spend, latency and GPU procurement. The third finding is cold water on the whole medical-AI-agent narrative: the most basic agentic competence is using a tool correctly, and right now even handing a model a calculator and getting the arithmetic right is unreliable.

Discount this

The v1.0 blog post is dated 2026-05-12, not this week. Sophont is itself a medical AI company, so building its own evaluation suite carries the same conflict of interest — the difference is that it open-sourced the items, the code and the leaderboard so third parties can reproduce and rebut it, which OpenEvidence did not. And the model versions on the board (Gemini 3 Pro Preview, GPT-5.2, GLM 4.7) have been superseded in the months since, so the current ordering may no longer hold.

Model efficiency Microsoft ResearchGigaPath-FlashApache-2.0

The one advance today that needs no scoreboard argument: GigaPath-Flash keeps 97% of pathology performance at 1/50th the compute — 22M + 21M parameters, Apache-2.0 weights

What

Microsoft Research released two efficient foundation models for computational pathology (arXiv 2607.18218v2). GigaPath-Flash pairs a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder, distilled from the billion-parameter GigaPath teacher, and retains 97% of GigaPath's average slide-level performance using 50x less compute; it averages 0.826 on PANDA and EBRAINS slide-level classification, beating tile-level-only baselines at 49.5x fewer FLOPs than the original. GigaTIME-Flash swaps the original GigaTIME CNN backbone for a ViT-S encoder, predicting 21-channel multiplex immunofluorescence maps from H&E images 6x faster with 8x less GPU memory: 14.9 GFLOPs per 256×256 tile against 69.1 (a 4.6x reduction), peak GPU memory of 2.16 GB at batch 128 against 16.68 GB, and throughput scaling to 1,679 tiles/second against the original's plateau near 390 — with higher mean Pearson correlation than the CNN original across both in-distribution and out-of-distribution cohorts. Both models' weights ship under Apache-2.0.

Why it matters

Because an efficiency gain does not need you to believe its scores — it lowers the barrier directly. The standing problem with pathology foundation models was never accuracy; it was that a whole-slide image runs to gigabytes and a machine that can process one is a capital expense, so only large academic centres played. Going from 16.68 GB to 2.16 GB is concrete: the first needs a data-centre GPU, the second runs on a consumer card. This is the only item today whose conclusion holds regardless of who grades it, and the one most likely to change who can actually run pathology AI over the next twelve months.

Discount this

All figures come from the authors' own paper (v2) and have not been independently reproduced; PANDA and EBRAINS are public research datasets, not prospective clinical validation — an average of 0.826 is not diagnostic-grade performance. The "97%" is a retention ratio on average slide-level performance and does not guarantee parity on every downstream task. This item is written from the arXiv HTML full text; the actual files and licence text on the weights pages were not separately verified.

Eval infrastructure MedAgentBenchStanfordNEJM AI

A picture that should not exist: on the most-cited medical agent benchmark, the top score is still Claude 3.5 Sonnet v2's 69.67% — and the board was last updated in May

What

MedAgentBench, built by a Stanford team, stands up a virtual EHR environment with 100 realistic patient charts holding 785,000 records and asks models to complete 300 clinical tasks, including ordering tests, prescribing and retrieving patient information (Becker's coverage; full study in NEJM AI). The trouble is that its public leaderboard was last updated on 2026-05-27, where the top entry remains Claude 3.5 Sonnet v2 at 69.67% (Query SR 85.33%, Action SR 54%). Note that 54% — the same model retrieves accurately, then gets it right half the time when it actually has to act.

Why it matters

This is the sharpest slice of medical AI's measurement problem: models turn over every few weeks while independent evaluation infrastructure lags by months. When a leaderboard indexed in NEJM AI and widely cited across the industry still has a model published in 2024 on top, anyone asking how strong the best medical agent is today has no option but to trust a vendor press release. The 85.33% Query SR against 54% Action SR deserves its own note — it points at the same thing as HealthAgentBench's 42%: there is a deep capability fault line between reading data and taking action, and it is the latter that hospitals buy.

Discount this

"Claude 3.5 Sonnet v2 still leads" means only that this third-party board has not been updated, not that no stronger model exists — it simply has not been independently tested and published, which is the point of the item. The Becker's coverage of the original study is dated 2025-09-16 and is old reporting; the leaderboard figures come from the third-party aggregator benchmarklist.com and were not reconciled line by line against the official MedAgentBench repository.

This week's preprints arXiv 2609.*9/29–9/30Harvard · BIDMC

What was actually new on 9/29–9/30: clinicians read through 19,930 conversations between young people and ChatGPT — the failure was not the wrong words but the right words at the wrong moment

What

The reading worth doing inside this week's window is all on arXiv. "Right Words, Wrong Moment" (2609.35953) had clinicians work through 19,930 conversations between young people and ChatGPT, and what they identified was not wrong content but process failure — most characteristically, prematurely jumping to solutions. ARCagent (2609.36392) tackles a very practical problem in clinical QA: how retrieval should calibrate when medical guidelines contradict each other, using ME/CFS as its case for "knowledge completeness and dynamic conflict-aware synthesis." MERID (2609.36238) uses recursively self-improving multimodal agents to combine interview and sensor data for major depression analysis. The same digests swept up UltraBench 2 (2609.28610, robust evaluation of vision foundation models on ultrasound) and nnFoundation (2609.26924, 3D foundation models for radiology). On the engineering side, Harvard/BIDMC's open-source digital psychiatry platform mindLAMP is now deployed across 17 countries and 65 sites on AWS serverless infrastructure.

Why it matters

The title "Right Words, Wrong Moment" is itself a one-line gloss on today's whole theme: benchmarks measure whether the content is right; clinical value turns on when it is said and in what order. A model scoring 100% on MedQA can still commit the process error of jumping to a solution too early in a real conversation, and no mainstream benchmark currently docks it for that. ARCagent's contradictory guidelines are the same issue from another side — real clinical difficulty is less often missing knowledge than choosing between pieces of knowledge that fight each other. As for UltraBench 2 and nnFoundation: when two independent groups propose new evaluation and new foundation models for ultrasound and 3D radiology in the same week, the imaging track is laying foundations, not raising ceilings.

Discount this

All of these are preprints without peer review. The authors, institutions and full data for arXiv 2609.35953, 2609.36392 and 2609.36238 could not be retrieved directly at the time of writing (fetches were refused); this item is written from the arXiv new-submissions listings and abstract-level descriptions, with scores and sample details not individually reconciled. The PDFs for 2609.28610 and 2609.26924 returned no machine-readable text and are listed on title and search-snippet only. The "17 countries, 65 sites" figure for mindLAMP comes from AWS's own industry blog and is vendor-reported.

02 — Product Analysis

Two ways of making a technical case: one rests on scores you announce yourself, the other on a compute bill anyone can re-run

OpenEvidence Model Family — Osler · Sackett · Snow · Darwin

A clinical QA model family tiered by reasoning depth · OpenEvidence (US)

Function and position. The design logic is reasoning depth matched to question difficulty: Osler answers in 5 seconds in the hallway, Sackett assembles consult-grade evidence in 30 seconds, Snow does a tumour-board-grade literature review in 5 minutes, all three free to US clinicians; Darwin is the research-preview ceiling model, withheld over dual-use risk (Becker's, 2026-09-08). The business model is free to physicians, monetized through the traffic — and what the $15 billion valuation on 9/30 bought is that distribution surface.

GigaPath-Flash / GigaTIME-Flash

Distilled computational pathology foundation models · Microsoft Research (US) · Apache-2.0

Function and position. It does not sell clinical interpretation; it sells making pathology AI runnable. GigaPath-Flash (a 22M ViT-S tile encoder plus a 21M LongNet slide encoder) distils the billion-parameter GigaPath into slide-level representations of whole-slide images; GigaTIME-Flash predicts 21-channel multiplex immunofluorescence maps of the tumour microenvironment from H&E. Both ship Apache-2.0 and are pretrained on large-scale real-world clinical data. The customer is not a hospital procurement office; it is everyone downstream building research or products on pathology images.

  • Strength : the efficiency figures can be re-run by anyone and need no trust in the vendor — 97% of performance at 50x less compute, 49.5x fewer FLOPs, 0.826 average on PANDA and EBRAINS, with GigaTIME-Flash cutting peak GPU memory from 16.68 GB to 2.16 GB and lifting throughput from about 390 to 1,679 tiles/second.
  • Strength : Apache-2.0 permits commercial use, modification and redistribution — not the half-open "research use only" posture. That decides whether it can actually be embedded in someone else's product.
  • Concern : PANDA and EBRAINS are research datasets, not prospective clinical validation; an average of 0.826 is some distance from diagnostic grade, and the "97%" is an average retention rate that does not guarantee parity on every downstream task. All figures come from the authors' own paper with no independent reproduction yet.

The contrast between the two is today's conclusion. OpenEvidence's case asks you to believe a set of scores it awarded itself, measuring multiple choice. GigaPath-Flash's case hands you a compute bill, and a compute bill cannot be faked — run it once and you know whether 2.16 GB is real. In a field where the yardstick itself is unreliable, the relative value of that second kind of progress — independently reproducible — is badly underrated.

03 — Companies & Competition

Who stands where, on what, against whom
Company Recent state & numbers Position & moat
OpenEvidence
ChatGPT for doctors, aiming to be the clinical OS
On 9/30 it raised $250M at a $15B valuation, about $1B over twelve months; on 9/3 it launched the four-model family, with Darwin at 100% on MedQA. The moat is physician-side distribution and habit, not the model — and the model side is squeezed from both directions by Nature Medicine's finding that general-purpose models win and by doubt over self-graded scores.
Microsoft Research (HLS)
Playing both hands: open weights plus independent benchmarks
GigaPath-Flash and GigaTIME-Flash shipped Apache-2.0 (97% of performance at 1/50th the compute), while HealthAgentBench's 54 tasks hold frontier agents to 42%. The moat is holding both the infrastructure and the scorecard: open-sourcing efficient models harvests an ecosystem, open-sourcing the benchmark lets it define what counts as progress. The weakness is that neither monetizes directly — both lean on Azure conversion.
Google / DeepMind
The other pole of open-weight medical models
MedGemma 1.5 (4B/27B) and MedASR: MRI findings classification at 65% (from 51%), chest X-ray anatomy localization IoU from 3% to 38%, EHRQA at 90% (from 68%); MedASR hits a 5.2% word error rate on chest X-ray dictations against Whisper large-v3's 12.5%, 58% fewer errors. Free for research and commercial use on Hugging Face and Vertex AI. The moat is the combination of free, commercially usable and small, which compresses the space for specialized medical model startups: when a 4B open-weight model handles CT, MRI, pathology and document understanding, selling a hospital a closed imaging model takes a lot more explaining.
Sophont
The medical AI startup that open-sourced its evaluation
Medmarks v1.0: 30 benchmarks, 61 models across 71 configurations, plus the Medmarks-T training environment; it finds GLM 4.7 needs over 5x the tokens to match GPT-5.2, and that giving models a calculator produced ignored tools, wrong answers and formatting failures. The moat is credibility rather than technology: in a market where vendor self-grading is the norm, "take my items and my code and run them yourself" is itself the differentiation. The weakness is that it also sells models, so its neutrality has a ceiling — and maintaining a 30-by-61 board takes sustained investment, with MedAgentBench frozen since May as the cautionary case.
Epic
The party that holds the data side
Its Comet predictive models are trained on over 100 billion data points from Cosmos de-identified records to forecast patient trajectories; it also announced several security programmes at this week's UGM. The moat is the data and the entry point to the workflow, not model quality — which is why HealthAgentBench's 42% matters more to Epic than to the model vendors: agents have to actually act inside the EHR before whose model is stronger becomes the relevant question.

Today's competitive structure is a squeeze. From above, general-purpose frontier models press down — Nature Medicine says they win, and Google's 4B open-weight medical model is free for commercial use. From below, efficient open models push up — an Apache-2.0 pathology foundation model runs at a fiftieth of the compute. Caught in between is exactly the closed, specialized, self-graded position — which is exactly the position that raised $250 million this week. The contradiction will not detonate immediately, because the valuation buys physician-side distribution rather than the model. But it explains why "nobody is grading this independently" is the objection being raised right now.

04 — Taiwan Angle

Taiwan built, two years ago, the thing the field is now missing most: a third-party validation centre

(1) The independent grading mechanism the field lacks already exists in Taiwan as institutional scaffolding. The Ministry of Health and Welfare launched three AI centres on October 7, 2024. The Clinical AI Evidence and Validation Centre works with the TFDA and pre-assembles a hospital consortium linked to electronic medical records, specifically to solve the difficulty of gathering validation data. The Responsible AI Implementation Centre requires public disclosure of models, data and performance. The AI Impact Research Centre works with the National Health Insurance Administration to build clinical trial mechanisms for assessing benefit, and from that a "scientific basis for pricing." At launch, 30 hospitals submitted 48 applications and 19 applications from 16 hospitals were selected, including NTUH, Taichung Veterans, Taipei Veterans and NCKU. Read against this week's dispute: what OpenEvidence lacks is not technology but precisely this kind of institution — someone outside the vendor doing the grading — and disclosing models, data and performance is the second centre's stated job. Having the institution is not the same as executing it, but this is a rare position where Taiwan leads the international conversation.

(2) GigaPath-Flash's compute bill means far more to Taiwan's mid-sized hospitals than to its medical centres. Peak GPU memory falling from 16.68 GB to 2.16 GB (arXiv 2607.18218v2) translates, in procurement terms, into the difference between a data-centre GPU and a single consumer card. Pathology AI deployment in Taiwan has long been stuck because regional and district hospitals have neither the compute nor the MLOps staffing, and an Apache-2.0 licence permits commercial use, modification and redistribution — meaning local vendors can build an on-premise deployable product from it rather than negotiating another licence. Paired with the TFDA's AI/ML medical device information and matchmaking platform, that route is considerably more realistic than waiting for a closed model to cut its price.

(3) The Ministry's generative AI guidance lands on exactly this week's gap. Taiwan has issued guidance for generative AI use by medical institutions, framing governance around six risk categories and nine key points, with detailed AI medical rules still being drafted. Set that against the finding in "Right Words, Wrong Moment" — that across 19,930 conversations the failures were process errors such as jumping to solutions too early, not wrong content. If the rules Taiwan writes govern only whether the output is correct, they will miss the category of error that actually harms patients. Governance documents should require that the interaction process be logged and auditable, not just that answers be accurate.

05 — Further Reading

These five form the full evidence chain for today's thread, ordered conclusion first, then method, then the rebuttal
  1. General-purpose chatbots outperform clinical AI tools on physicians' real-world questions — Nature Medicine research briefing (2026-06-17)

    The two-page briefing reads better than the paper, and the line about specialized tools doing no better than a Google AI overview is the anchor for every argument today.

  2. Medical AI has a measurement problem — Nature, Arjun K. Manrai, Harvard Medical School (2026-07-28)

    It turns the unreliable-yardstick problem into something you can reason about rather than complain about; after it you see why the self-grading dispute is methodological, not a PR matter.

  3. HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments — arXiv:2606.31179 (2026-06-30)

    If you want the real ceiling on medical AI agents right now, how these 54 tasks are constructed tells you more than any single score.

  4. Release of Medmarks v1.0, technical report, and updated leaderboard — Sophont Blog (2026-05-12)

    The only report that writes a deployment cost like "open models burn 5x the tokens" into its evaluation conclusions — more useful than the ranking if you are the one choosing a model.

  5. OpenEvidence, TEMPO & No One's Watching — Grid Health (2026-09-22)

    The cleanest statement of the other side, and it frames the FDA's TEMPO pilot (expanding from four participants to forty) alongside vendor self-grading — both substituting real-world deployment for independent pre-market verification.

06 — References

References
  1. News 9/30/26. HISTalk, 2026-09-30. histalk2.com
  2. OpenEvidence, TEMPO & No One's Watching. Grid Health, 2026-09-22. gridhealth.io
  3. OpenEvidence's new model achieves 1st perfect score on medical AI benchmark. Becker's Hospital Review, 2026-09-08. beckershospitalreview.com
  4. Introducing the OpenEvidence Model Family. Business Wire, 2026-09-03. businesswire.com
  5. OpenEvidence launches 4 medical AI models for clinicians. MobiHealthNews. mobihealthnews.com
  6. General-purpose chatbots outperform clinical AI tools on physicians' real-world questions. Nature Medicine 32(7):2364–2365, 2026-06-17. DOI 10.1038/s41591-026-04457-9. nature.com
  7. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine, 2026. nature.com · PubMed 42286322
  8. Manrai AK. Medical AI has a measurement problem. Nature, 2026-07-28. nature.com
  9. Liu Q, Zhang S, et al. (Poon H, corresponding). HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents. arXiv:2606.31179, 2026-06-30. arxiv.org
  10. Release of Medmarks v1.0, technical report, and updated leaderboard. Sophont Blog, 2026-05-12. sophont.med
  11. Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks. arXiv:2605.01417. arxiv.org
  12. GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis. arXiv:2607.18218v2, Microsoft Research. arxiv.org
  13. MedAgentBench Benchmark Scores & AI Model Leaderboard. benchmarklist.com, last updated 2026-05-27. benchmarklist.com
  14. MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents. NEJM AI, DOI 10.1056/AIdbp2500144. ai.nejm.org
  15. Stanford benchmarks AI agents in healthcare. Becker's Hospital Review, 2025-09-16. beckershospitalreview.com
  16. Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT. arXiv:2609.35953, 2026-09. arxiv.org
  17. ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering. arXiv:2609.36392, 2026-09. arxiv.org
  18. MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis. arXiv:2609.36238, 2026-09. arxiv.org
  19. UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound. arXiv:2609.28610, 2026-09. arxiv.org
  20. nnFoundation: 3D Foundation Models for Radiology. arXiv:2609.26924v1, 2026-09. arxiv.org
  21. 醫療 AI 行業日報 2026-09-30 / Medical AI industry daily digest. agents-radar, Issue #343. github.com
  22. 醫療 AI 行業日報 2026-09-29 / Medical AI industry daily digest. agents-radar, Issue #334. github.com
  23. From cloud to clinic: How AWS powers digital psychiatry at scale. AWS Industries Blog. aws.amazon.com
  24. Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR. Google Research Blog, 2026-01-13. research.google
  25. google/medgemma-1.5-4b-it. Hugging Face. huggingface.co
  26. Epic Launches Comet: A New AI Platform to Predict Patient Health Journeys. HIT Consultant, 2025-09-03. hitconsultant.net
  27. 衛生福利部三大AI中心啟動記者會 / MOHW launches three AI centres. 衛生福利部, 2024-10-07. mohw.gov.tw
  28. 衛福部首發醫療GenAI指引 六大風險、九項要點建立治理框架. 環球生技月刊 / GBI Monthly. news.gbimonthly.com
  29. 厚生會成立智慧醫療委員會 衛福部擬提AI醫療細則草案. 中央社 / CNA, 2026-08-06. cna.com.tw
  30. 智慧醫療器材資訊暨媒合平台 / AI-ML Medical Device Information and Matchmaking Platform. TFDA. aimd.fda.gov.tw
  31. Insurers claim AI is already increasing healthcare costs. TechCrunch, 2026-09-26. techcrunch.com
Editor's note: (1) New publications in this theme were thin today. Within the 9/28–10/1 window no flagship medical model shipped; what actually landed were arXiv preprints and Hugging Face weights. This issue therefore enters through this week's funding round and self-grading dispute, and pulls the key evaluation studies from May–July 2026 back in as the yardstick (Nature Medicine research briefing 6/17, Nature commentary 7/28, HealthAgentBench 6/30, Medmarks v1.0 5/12, GigaPath-Flash v2 in July), flagging each original publication date in that item's "Discount this" row rather than passing older work off as news. (2) The Nature Medicine full text at s41591-026-04431-5 returned a 503 at the time of writing; that item is built from the research briefing (DOI 10.1038/s41591-026-04457-9) and the PubMed 42286322 record, not the full text. The briefing names neither the three general-purpose models nor the two specialized tools, and publishes no per-model scores. (3) Fetches for arXiv 2609.35953, 2609.36392 and 2609.36238 were refused, and the PDFs for 2609.28610 and 2609.26924 returned no machine-readable text; all five are written from arXiv new-submission listings and abstract-level descriptions, with authors, institutions and full data not individually reconciled. (4) OpenEvidence's four benchmark scores, mindLAMP's "17 countries, 65 sites" and Epic Comet's "100 billion-plus data points" are all vendor-reported and not independently audited. (5) Parts of Becker's and HISTalk sit behind a paywall or newsletter; only publicly visible passages were used. The round size and valuation come from HISTalk's secondary coverage, which itself notes a discrepancy with earlier reporting. (6) Downloads for Fastino-Nemotron-3.5-Lightning-Healthcare were reported as 6,825 on 9/29 and 6,707 on 9/30, both from the agents-radar digest and unverified against Hugging Face; the 9/30 figure is used above. (7) Model versions on the Medmarks leaderboard and the top entry on MedAgentBench may have been superseded in the months since publication; the rankings reflect only the state at each board's last update.