◆ AI & Medical AI Daily
–
Sunday · Weekly Review & Further Reading

This was the week the claim that AI outperforms doctors hit its peak — and the week every piece of evidence pulled it back down: McKinsey put clinical AI at 16–22% of US outpatient care on 28 Sept (2–3bn claims a year), HHS Secretary Kennedy told the MAHA Summit on 29 Sept that an AI second opinion is “much better informed than any doctor in the country”, six physician organisations rebutted him on 30 Sept, and the hospital CIOs actually asked put the real number at “1–2% of primary and urgent care”; the one autonomous AI that shipped third-party validation this week does billing

The most useful thing this week produced was two answers to one question. On 28 September McKinsey worked through 2024 Merative commercial claims plus the CMS Medicare Limited Data Set and the Medicaid VRDC and concluded that clinical AI “can perform some 16 to 22 percent of US outpatient care” — roughly 2–3bn claims a year, 13–19% of outpatient spending. Three days later Becker's put the same report to eight hospital CIOs and informatics chiefs, and Parkview Health's CMIO Mark Mabus answered: “I am very doubtful that we are actually already at 22% … My generous estimate is more like 1-2% of primary care/urgent care only.” The two figures are more than an order of magnitude apart, and the gap does not sit in model capability — it sits in who will use the thing and who answers for it. Everything else this week — Kennedy's “medical tyranny”, the six societies' joint statement, AKASA coding an inpatient stay in 90 seconds, Oura shelving a $2.2bn IPO, Anthropic reportedly listing in November at a $2tn valuation — reads better placed inside that tenfold gap.

01 — Top Stories

Seven things, ordered claim → rebuttal → measurement → capital
Market data McKinseyCMS / Merative9/28

McKinsey turns “how much outpatient care can AI do” into a number from billing codes: 16–22%, 2–3bn claims

What

“The AI-powered future of care”, published 28 September, finds clinical AI can already perform 16–22% of US outpatient care — some 2–3bn claims a year, or 13–19% of outpatient spending. It splits in two: evaluation and management visits at 11–15% of outpatient claims (1.5–2.0bn), and diagnostic and imaging interpretation at 5–7% (700–900m). On labour, AI can execute substantial parts of the core tasks of roughly 11 million of the 19 million people working in US health care, close to 60%. The method: 2024 Merative commercial claims plus the CMS Medicare Limited Data Set and the Medicaid Virtual Research Data Center, automatable claims identified by billing code and procedure modifier, then task-level exposure mapped through O*NET occupational data (McKinsey, 2026-09-28; Becker's summary).

Why it matters

It became the common citation behind every argument this week, and it is useful precisely because it does not measure whether AI is accurate — it measures how much work is technically takeable. Billing codes are off-the-shelf and auditable, and no model grades itself. But it does not measure how much has actually been taken, and the industry spent the next three days prying that distinction open.

Discount this

This is a consultancy's ceiling on technical feasibility, not a deployment rate. It sits on 2024 claims data, and two years is not a short lag in this field. And “AI-performable” is inferred by sorting billing codes, while patient complexity inside a single E&M code varies enormously.

Policy HHSMAHA Summit9/29

HHS Secretary Kennedy: an AI second opinion is “much better informed than any doctor in the country” and frees Americans from “medical tyranny”

What

On 29 September, in a fireside chat with Vice-President JD Vance at the Make America Healthy Again Summit in Washington, HHS Secretary Robert F. Kennedy Jr. said AI can give patients a second opinion “much better informed than any doctor in the country”, that it would free them from “medical tyranny”, and that it would let Americans independently check official advice on masks, social distancing and vaccines. He also said the administration would ensure “every American will have access on their cell phones to their own medical records”, and that AI could read the long records clinicians have no time for. He relayed OpenAI chief executive Sam Altman's line that it would be malpractice for a physician to diagnose or prescribe “without at least checking AI” (Fierce Healthcare, 2026-10-02; Medical Economics).

Why it matters

This is not a vendor's marketing line; it is the head of the US health department publicly positioning AI as a tool against professional authority. Two concrete things turn on it: how hard record portability gets pushed, which most of the profession supports, and which way the FDA's generative-AI device framework lands — comments on that discussion paper close on 19 October (FDA).

Discount this

“Better informed than any doctor in the country” came with no study attached. The head-to-head evidence in circulation this week points elsewhere anyway: the Nature Medicine paper from last week found general-purpose frontier models beat two specialised clinical AI tools on medical benchmarks — but that is model against model, not model against physician (Nature Medicine).

Profession AMAAAFP / AAP / ACOG / ACP / ACS9/30

Six physician bodies reply the next day: statements implying AI is inherently better informed “diminish physician expertise and risk undermining patients' trust”

What

On 30 September the American Medical Association, American Academy of Family Physicians, American Academy of Pediatrics, American College of Obstetricians and Gynecologists, American College of Physicians and American College of Surgeons issued a joint statement. It grants that “AI has tremendous potential to provide new tools and insights” but rejects framing that pits physicians against AI. The load-bearing sentence: statements suggesting AI is inherently better informed than physicians “diminish physician expertise and risk undermining patients' trust”. It adds that “physicians evaluate patients in context” and that “patient safety, physician expertise and the humanity of clinical practice must guide how these tools are developed” (AMA press release, 2026-09-30; ACP).

Why it matters

Six societies co-signing inside 24 hours is a rare move in American medicine, and the term they chose is augmented intelligence rather than artificial intelligence. That word fight is really a liability fight: if it is augmented, the decision stays with the physician; if it is artificial and can independently issue a second opinion, someone has to be reassigned the responsibility. Nobody answered that this week.

Reality check Parkview / Stanford / UCIBecker's10/2

Eight hospital CIOs mark 22% down to 1–2%: “I have absolutely nothing nice to say about Secretary Kennedy's opinion”

What

Becker's put the McKinsey number and Kennedy's claim to eight hospital IT and informatics leaders. Parkview Health CMIO Dr Mark Mabus: “I am very doubtful that we are actually already at 22% … My generous estimate is more like 1-2% of primary care/urgent care only.” TidalHealth CIO Dr Mark Weisman thought the figure realistic for primary and urgent care. Stanford Health Care chief data scientist Dr Nigam Shah noted it likely does not apply at tertiary centres, where low-complexity outpatient work is not the largest volume. UCI Health CMIO Dr Deepti Pandita: “The McKinsey estimate is directionally plausible, but the reality is far more nuanced.” On Kennedy, Monument Health CIO Dr Patrick Woodard: “I have absolutely nothing nice to say about Secretary Kennedy's opinion.” Reid Health CIO Muhammad Siddiqui said the claim “goes further than the evidence supports”. Hospital for Special Surgery chief transformation officer Dr Ashis Barad landed it on liability: “Someone has to decide what good care looks like … and take responsibility when a patient acts on its advice.” (Becker's Hospital Review, 2026-10-02)

Why it matters

These answers show the tenfold gap is not a disagreement but two different rulers: McKinsey measures work that can be taken, the CIOs measure work that has been taken. More telling, none of them disputes technical feasibility — Weisman endorses it outright — and every objection lands on liability, context and institutional complexity. Which is exactly the shape of the Rambam emergency-department study this daily covered on 28 September: 99 of 100 sampled LLM outputs rated clinically appropriate, while physician uptake fell from 68% to 30% over four weeks (Nature Medicine, 2026-08-19).

Product AKASACleveland Clinic10/2

The only autonomous AI to ship third-party validation this week does billing: AKASA cuts inpatient coding from 30–60 minutes to under 90 seconds

What

On 2 October AKASA commercially launched an autonomous mid-cycle AI platform for inpatient coding and clinical documentation integrity, running on complex hospital charts without human intervention. The company says it completes inpatient coding in under 90 seconds post-discharge against the 30–60 minutes a human coder typically takes; it currently handles roughly 1 in 10 US inpatient discharges, with inpatient volume up nearly 6x year on year. Its client base represents over $180bn in aggregate net patient revenue, with Cleveland Clinic and Nebraska Methodist Health System named. Third-party validation showed the platform “matched or exceeded the accuracy of experienced human medical coders” on MS-DRG assignment, principal diagnosis selection, present-on-admission indicators and quality documentation capture (HIT Consultant, 2026-10-02; Fierce Healthcare).

Why it matters

Lay this on top of McKinsey's 16–22% and the week's real conclusion comes into focus: autonomous AI is already running at scale in American hospitals, but it runs in the revenue cycle, not in diagnosis. The reason is not mysterious. Billing has a determinate right answer — an MS-DRG code is correct or it is not — an existing audit apparatus, and failures that cost money rather than people. What hospitals are replacing with it is manual BPO outsourcing, not physicians. It is the same logic that makes Abridge's VA award (22 September, a five-year enterprise IDIQ with a $775.72m ceiling across 1,380 facilities and 170 VA medical centres) a documentation contract rather than a decision-making one (HIT Consultant).

Discount this

The 90 seconds, the 6x, the $180bn and the “1 in 10 inpatient discharges” are all vendor-reported. Third-party validation is mentioned but the validator is unnamed, with no published method or sample size and no peer review. If autonomous coding systematically over-assigns DRGs, the consequence is CMS audit and recoupment — a risk that normally only surfaces several audit cycles later.

Capital AnthropicARPA-H9/29 · 10/1

Anthropic did two things in one week: a reported November IPO at up to a $2tn valuation, and joining ARPA-H's clinical AI moonshot

What

On 1 October Bloomberg reported Anthropic targeting a mid-November IPO, with formal marketing possibly starting the week of 9 November and a valuation that could reach or exceed $2tn; selected institutional investors have been invited to its San Francisco headquarters on 14 October (Bloomberg / Yahoo Finance, 2026-10-01; Briefs). Two days earlier STAT reported Anthropic joining ARPA-H's clinical AI moonshot and planning a closed-door health care event. That programme is ADVOCATE, launched 9 September at $62.7m over four years as the world's first bid to build FDA-authorised clinical AI in cardiovascular care, with Atman Health, Tempus AI and Updoc already in (STAT, 2026-09-29; ARPA-H).

Why it matters

For two years the frontier labs' health strategy has largely been to sell models to health companies. ADVOCATE forces a different route: participants have to carry a product through to FDA authorisation, which means taking regulatory responsibility for a specific clinical claim. A frontier lab heading into a listing agreeing to stand there suggests the general-model-versus-specialised-tool fight is moving off the benchmark and into the regulatory file — which is the part of that fight benchmarks cannot win.

Discount this

The IPO timing and valuation come entirely from press reports citing unnamed sources; Anthropic has not confirmed publicly, and $2tn is a could-reach ceiling, not a price. The STAT item sits behind a paywall and is written up here only from the publicly visible headline and lede: Anthropic's specific role, contribution and division of work inside ADVOCATE are not publicly disclosed.

Technology StanfordClaude10/1

Stanford's chair of medicine: Claude read his whole genome in 30 minutes for about $5 — the same job took 30 people nearly a year in 2009

What

On 1 October STAT published an opinion piece by Euan Ashley, chair of the department of medicine at Stanford. Using a 180-word prompt and roughly 400,000 tokens, he asked Claude to analyse his own whole genome under the same framework his team built in 2009 — identify rare disease variants, pharmacogenomic variants and assess common disease risk. It took 30 minutes and about $5, against 30 people and nearly a year in 2009. It correctly found that he carries two APOE ε4 variants associated with Alzheimer's risk, matching the original analysis. He calls for well-characterised reference genomes reflecting global population diversity, a shared catalogue of medically significant genes that are technically hard to sequence, clear minimum detection thresholds tailored to the application, and transparent disclosure from both labs and AI developers about what they reliably detect (STAT, 2026-10-01).

Why it matters

This is the week's single record of AI actually doing what used to need a team, and the author led that team, so the baseline is credible. What matters is that his conclusion is not “therefore ship it” but “therefore we urgently need standards” — all four of his asks are about disclosing the boundary of detection, which is precisely what Kennedy's claim lacked. An output that correctly names APOE ε4 does not thereby tell you where it failed to look.

Discount this

n = 1, the subject is the author, and he knew the 2009 answers going in — this is a demonstration, not an evaluation, and it cannot bound a false-negative rate. The piece sits behind STAT's paywall and is written up here from the publicly visible content.

Markets OuraHeidi / Tiny Health / Grindr9/29 · 10/1

One exit window shut this week: Oura shelves a $2.2bn IPO — while the private market still wrote a $340m cheque

What

On 29 September smart-ring maker Oura shelved its US IPO — 55m shares at $40–44, up to $2.2bn, implying as much as $15bn at the midpoint. Its 2025 revenue was $907.9m, with 2026 expected up 90%. The company cited only “uncertainty in the IPO market”; chief executive Tom Hale said “we have the luxury of choosing our moment” (TechCrunch, 2026-09-29; CNBC). The private market did not cool in step: Heidi closed a $340m Series C at a $900m valuation on 22 September for agentic clinical workflow automation; Tiny Health closed a $33m Series B on 29 September (taking it to $46m); Basalt Health took $20m in a Series A on 24 September. On the M&A side, Grindr announced on 1 October that it is buying PurposeMed for $250m in cash and stock with up to $70m in further consideration, closing in Q4 (Health IT Answers, 2026-10-01; Fierce Healthcare weekly rundown).

Why it matters

A consumer health hardware company with close to $1bn of revenue growing 90% chose not to list, in the same week Anthropic was reported heading to market at $2tn. The public market is pricing “AI infrastructure” and “AI applied to health” by completely different machinery. For Taiwanese medical AI founders the readable signal is not to plan an exit around a US listing: the private round and the strategic acquisition are the doors still open.

Discount this

“Market uncertainty” is the standard formula for a shelved listing; the company did not say whether the real reason was soft demand, a valuation gap or simply waiting on another quarter, and it should not be over-read. The Heidi and Basalt rounds date to 22 and 24 September, strictly before this week's window, and are included here only as a read on concurrent private-market temperature.

02 — Product Analysis

Two autonomous systems actually running this week — one on billing, one on genomes — separated by who answers for a mistake

AKASA Autonomous Mid-Cycle AI

No-touch inpatient coding and CDI · AKASA (US)

Function and position. It reads the full inpatient chart and autonomously produces MS-DRG assignment, principal diagnosis selection, present-on-admission indicators and CDI queries. The buyer is the revenue-cycle department of a large academic or regional system, and what it displaces is manual BPO outsourcing, not clinical staff. Commercially launched 2 October, claiming completion within 90 seconds of discharge against 30–60 minutes for a human coder (HIT Consultant).

  • Strength : this is the only product this week that puts “autonomous” on a task with a determinate right answer. MS-DRG codes can be audited after the fact and errors quantified, which is why it can make a claim like third-party validation “matched or exceeded” experienced human coders — a sentence that cannot even be formed in a diagnostic setting (Fierce Healthcare).
  • Strength : scale is already the moat. It claims roughly 1 in 10 US inpatient discharges, inpatient volume up nearly 6x year on year, and a client base representing over $180bn in net patient revenue including Cleveland Clinic. A coding model improves directly with the DRG distribution it has seen and the audit feedback it has absorbed, which is the hardest thing for an entrant to replicate (HIT Consultant).
  • Concern : the validator is unnamed, with no published method or sample size and no peer review, and every performance figure is vendor-reported. The more structural risk is the direction of error: if autonomous coding drifts systematically toward higher-paying DRGs, that surfaces as CMS recoupment and potential False Claims Act exposure several audit cycles later, not on today's accuracy sheet.
  • Concern : the pricing model is undisclosed. The economics of revenue-cycle automation turn on whether it is per-chart, per-cent-of-collections or per-seat, and on whether the vendor shares in incremental revenue — the last of which has long been a regulatorily sensitive structure in the US. Neither the company nor the coverage addresses it.

Claude as a whole-genome interpreter

A general frontier model as a genome interpretation pipeline · Anthropic (US) / Stanford demonstration

Function and position. Not a product but a demonstration of a general model being used as one. Stanford's chair of medicine Euan Ashley used a 180-word prompt and about 400,000 tokens to have Claude interpret his own whole genome under his 2009 framework — 30 minutes, about $5, correctly identifying his two APOE ε4 variants (STAT, 2026-10-01).

  • Strength : the cost shift is the largest single number in the week. Thirty people over nearly a year becomes 30 minutes and $5 — a drop of that magnitude is not an efficiency gain, it changes who is able to do the work at all. And the baseline comes from the person who ran the original team, which is more credible than any vendor self-test (STAT).
  • Strength : having no moat is itself an advantage here — the route needs no genomics-specific model and no licensed dataset, just a prompt and a large enough context window. For a Taiwanese medical centre this is one of the few capabilities reproducible without first negotiating a procurement contract.
  • Concern : n = 1, the subject is the author, and he knew the 2009 answers going in. The demonstration can establish what is detectable and says nothing at all about what is not — and clinical genome interpretation carries nearly all its risk in false negatives and unreported pathogenic variants. Three of Ashley's own four asks (reference genomes, a catalogue of hard-to-sequence genes, minimum detection thresholds) exist to plug exactly that hole.
  • Concern : ancestry skew in the reference data is the unquantified part. Ashley explicitly asks for reference genomes “reflecting global population diversity”, which is an admission that performance in non-European ancestries is currently unknown. For a medical centre serving an East Asian population, that is something to measure locally before quoting the $5 figure.

03 — Companies & Competition

Who stands where, on what, against whom
Company Recent state & numbers Position & moat
AKASA
Autonomous inpatient coding and CDI
Launched its autonomous inpatient platform on 2 Oct, claiming coding inside 90 seconds post-discharge, roughly 1 in 10 US inpatient discharges, inpatient volume up nearly 6x year on year, and clients representing over $180bn in net patient revenue (Cleveland Clinic, Nebraska Methodist) (HIT Consultant, 2026-10-02). The moat is accumulated audit feedback plus academic-centre references. The thin spot: it competes on the price of outsourced labour, and if Epic builds autonomous coding natively its distribution advantage would immediately outweigh any model advantage.
Abridge
Ambient clinical documentation
Selected on 22 Sept for the VA enterprise contract — a five-year multiple-award IDIQ with a $775.72m ceiling, alongside Knowtex — covering 9m+ veterans across 1,380 facilities (170 VA medical centres, nearly 1,200 outpatient sites), already live at 75+ VA medical centres post-pilot, spanning both VistA/CPRS and the Oracle Health federal EHR, with multilingual support validated in 28+ languages (HIT Consultant; Nextgov). The moat is working across two mutually incompatible federal EHRs — engineering debt a pure model company struggles to take on. The limit: it still sells documentation rather than decisions, so the ceiling is ambient-scribe ARPU. And the headline figure is a shared multiple-award ceiling, not Abridge's order book.
Anthropic
Frontier models, turning toward the regulatory path
Reported on 1 Oct to be targeting a mid-November IPO at a valuation that could reach or exceed $2tn, with institutional investors invited to San Francisco on 14 Oct; reported on 29 Sept to have joined ARPA-H's clinical AI moonshot (ADVOCATE, $62.7m over four years) (Bloomberg / Yahoo; STAT). Its health moat is capability spillover — Ashley's genome demonstration is a free capability proof — plus a new willingness to walk the FDA authorisation path. The weakness: it owns neither clinical data nor the workflow entry point, and its role and contribution in ADVOCATE are not publicly disclosed.
Oura
Consumer wearables and health sensing
Shelved an IPO of up to $2.2bn on 29 Sept (55m shares at $40–44, implying up to $15bn at the midpoint). 2025 revenue of $907.9m with 2026 expected up 90%. The only stated reason was “uncertainty in the IPO market”; chief executive Tom Hale: “we have the luxury of choosing our moment” (TechCrunch, 2026-09-29). The moat is the ring form factor and longitudinal physiological data, against Apple, Samsung and Whoop. The weakness is that it is still priced as consumer electronics rather than as health care — and this week's shelving is in substance the result of that pricing problem.
Heidi
Agentic clinical workflow
Closed a $340m Series C at a $900m valuation on 22 Sept, funding “AI-powered workflow automation and agentic capabilities for clinicians” (Health IT Answers, 2026-10-01). It is walking upstream from scribe to agent, straight into Abridge and Microsoft/Nuance. The weakness: $340m against a $900m valuation is an unusually high ratio, which implies the money is to buy distribution rather than a technical lead.
Welldoc
Chronic care loop and patient-side data
Launched an AI-driven care loop platform on 1 Oct connecting patients' daily health data back to care teams, which it says is powered by 500m+ health data points across 80+ health domains (Fierce Healthcare weekly rundown, 2026-10-02). The moat is the time depth of longitudinal chronic-disease data. The weakness is that “500m data points” measures size, not effect: the launch carried no outcome data at all, which is precisely the kind of evidence Nature Medicine's editors have been demanding.

This week's competitive structure is a layering: where there is a determinate answer, autonomy has already won; where there is not, the fight is still over who gets to define the answer. AKASA and Abridge are negotiating contract scale and distribution in the revenue cycle and the documentation layer — a $775.72m IDIQ ceiling, 1 in 10 inpatient discharges — while everything that happened in the diagnostic layer this week was rhetoric: the Secretary's claim, six societies' rebuttal, eight CIOs' measured gap. Anthropic walking into ADVOCATE is the single exception to that layering: it chose not to win another benchmark and to carry an FDA-authorised clinical claim instead.

04 — Taiwan Angle

America spent the week arguing whether AI can replace doctors; Taiwan next week is still arguing whether the data can be joined up

(1) The timing lines up: what Taiwan's health minister is due to say on 7 October is exactly what this week's American argument was missing. Economic Daily News holds its 2026 Biotech Forum on 7 October, where Minister of Health and Welfare Shih Chung-liang will speak, arguing that for medical AI to land in practice “the first task is to establish data standards, governance systems and a trustworthy environment for use”. Concrete milestones include adopting the FHIR health information exchange standard, connecting core clinical data across all of Taiwan's medical centres by year-end, and extending that to all hospitals by the end of next year. On budget: NT$48.9bn over five years for the Healthy Taiwan Deep Cultivation Plan, NT$24bn over four years for the National Drug Resilience Preparedness Plan, and a planned NT$10bn in strategic investment (Economic Daily News, 2026-09-15). Lay that against the eight CIOs in Becker's this week: not one of their objections to Kennedy was that the models are not good enough. Every one landed on context, liability and institutional complexity — which is to say, on governance.

(2) McKinsey's 16–22% cannot be lifted into Taiwan, because the number is built out of American billing codes. The method rests on classifying billing codes and procedure modifiers in Merative commercial claims, the CMS Medicare Limited Data Set and the Medicaid VRDC (McKinsey). Taiwan has a single payer whose fee-schedule code structure does not follow US E&M levels, and both outpatient volume and the time density of a visit differ substantially. So 16–22% is neither a ceiling nor a floor here — it is a calculation that has to be redone. What does transfer is the method: run the same classification over National Health Insurance claims and you would learn where Taiwan's automatable outpatient share actually sits. On the current data-governance timetable, that only becomes possible once the medical-centre data linkage completes at year-end.

(3) The thing America's six physician organisations did this week, Taiwan's equivalent bodies have not done. The AMA and five others co-signed a statement within 24 hours, substituting “augmented intelligence” for “artificial intelligence” in order to fix where responsibility sits (AMA). Taiwan's institutional base is not thin: the TFDA established its Smart Medical Device Project Office on 7 May 2021 as a single-window consulting and regulatory-guidance channel for AI/ML devices (MOHW) and runs an AI/ML medical device information and matchmaking platform (aimd.fda.gov.tw); the three AI centres launched on 7 October 2024 — Responsible AI Implementation, Clinical AI Evidence Verification, and AI Impact Research — cover transparency and reliability, certification evidence, and the basis for insurance reimbursement respectively, with 16 hospitals and 19 proposals approved (MOHW). What is missing is a collective position from the profession itself: if a Kennedy-style claim is made here, there is currently no document co-signed by physician bodies defining who answers for an AI output. This week's American statement is worth reading as a template.

05 — Further Reading

Chosen so that reading it changes your judgement of this week's argument rather than adding detail. The first three are this week's primary documents; the last three are the methodology you calibrate them against.
  1. The AI-powered future of care — McKinsey & Company (2026-09-28)

    The common citation behind every argument this week, and its methodology section is worth more than its conclusion: it states exactly how “AI-performable” is derived from billing codes and procedure modifiers. Read that and you can discount the 16–22% yourself.

  2. AI as patients' 1st stop for care? 8 health system leaders on RFK Jr., McKinsey — Becker's Hospital Review (2026-10-02)

    Eight named, titled, direct quotes — the most honest document of the week. The point is not what they object to but that none of the objections is about accuracy. Read as a list, it is where medical AI procurement will get stuck for the next three years.

  3. Statement from leading physician organizations on the role of augmented intelligence in healthcare — AMA et al. (2026-09-30)

    Short, but every word was negotiated. Note that it says “augmented intelligence” throughout and never “artificial intelligence”: the substitution is not fastidiousness, it is drawing the liability line. If Taiwanese physician bodies draft an equivalent, this is the wording to work from.

  4. Prospective evaluation of a large language model clinical decision support system in the emergency department — Nature Medicine (2026-08-19)

    The answer to this week's puzzle is in here. Rambam's emergency department, 1,138 patients, four weeks, two parallel units: 99 of 100 sampled outputs rated clinically appropriate, while physician uptake fell from 68% to 30% and dropped with every extra hour into a shift (OR = 0.72 per shift hour); length of stay was 4.9 h in both arms (P = 0.99). The authors conclude the barrier is sustained clinician engagement, not algorithmic accuracy.

  5. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks — Nature Medicine

    Read it to see what “better than doctors” turns into when someone actually runs a head-to-head: general frontier models against specialised clinical AI tools, with the specialised tools losing across MedQA, HealthBench and real physician questions. Note it is model against model, so it neither supports nor refutes Kennedy — which is the point, because nobody has run the study he would need.

  6. Show us the evidence for the value of medical AI — Nature Medicine editorial (2026-04-21)

    An April editorial, but this is the week to reread it. Its argument is that evidentiary standards should scale with the strength of the claim, and that technical performance metrics alone cannot carry a clinical value claim. Hold that ruler against Kennedy's “better informed than any doctor in the country” and Welldoc's “500m data points” and neither measures up.

  7. Claude analyzed my genome in 30 minutes. Now we need standards for the results — Euan Ashley, STAT Opinion (2026-10-01)

    What to read is not the 30-minutes-and-$5 comparison but the four standards in the second half. Someone who has just demonstrated the capability himself immediately asks only for disclosure of the detection boundary — the week's clearest demonstration of the distance between capability and deployability. Paywalled; read alongside this report's summary of the publicly visible portion.

06 — References

References
  1. The AI-powered future of care. McKinsey & Company, 2026-09-28. mckinsey.com
  2. AI can now perform up to 22% of US outpatient care: McKinsey. Becker's Hospital Review, 2026-09-29. beckershospitalreview.com
  3. AI as patients' 1st stop for care? 8 health system leaders on RFK Jr., McKinsey. Becker's Hospital Review, 2026-10-02. beckershospitalreview.com
  4. RFK Jr.'s AI claims spark physician backlash as debate over medical AI intensifies. Fierce Healthcare, 2026-10-02. fiercehealthcare.com
  5. Six physician groups push back on claims that AI is better informed than doctors. Medical Economics, 2026-10. medicaleconomics.com
  6. Statement from leading physician organizations on the role of augmented intelligence in healthcare. American Medical Association, 2026-09-30. ama-assn.org
  7. Statement from leading physician organizations on the role of augmented intelligence in healthcare. American College of Physicians, 2026-09-30. acponline.org
  8. AKASA Launches Autonomous Mid-Cycle AI for Inpatient Coding and CDI. HIT Consultant, 2026-10-02. hitconsultant.net
  9. AKASA debuts autonomous AI platform for inpatient medical coding and clinical documentation. Fierce Healthcare, 2026-10-02. fiercehealthcare.com
  10. Abridge Wins Seat on $775.7M VA Enterprise Contract to Power Ambient Clinical AI. HIT Consultant, 2026-09-22. hitconsultant.net
  11. VA selects Abridge ambient scribe under new enterprise contract. Nextgov/FCW, 2026-09. nextgov.com
  12. Anthropic targets pre-Thanksgiving IPO at $2 trillion valuation. Bloomberg via Yahoo Finance, 2026-10-01. finance.yahoo.com
  13. Anthropic Sets Investor Meetings on Oct. 14 Ahead of Possible IPO. Briefs, 2026-10. briefs.co
  14. Anthropic joins ARPA-H clinical AI moonshot, will hold closed-door health care event. STAT, 2026-09-29(付費牆/paywalled). statnews.com
  15. ARPA-H launches world's first bid to build FDA-authorized clinical AI (ADVOCATE). ARPA-H, 2026-09-09. arpa-h.gov
  16. Claude analyzed my genome in 30 minutes. Now we need standards for the results. STAT Opinion, 2026-10-01(付費牆/paywalled). statnews.com
  17. Oura shelves its $2.2B IPO, citing 'uncertainty' in the market. TechCrunch, 2026-09-29. techcrunch.com
  18. Smart ring maker Oura postpones IPO due to market 'uncertainty'. CNBC, 2026-09-29. cnbc.com
  19. Health IT Business News, Financial Edition. Health IT Answers, 2026-10-01. healthitanswers.net
  20. Weekly Rundown: Duke University launches digital twin program; Welldoc unveils AI-powered care loop. Fierce Healthcare, 2026-10-02. fiercehealthcare.com
  21. Prospective evaluation of a large language model clinical decision support system in the emergency department. Nature Medicine, 2026-08-19. nature.com
  22. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine, 2026. nature.com
  23. Show us the evidence for the value of medical AI. Nature Medicine editorial, 2026-04-21. nature.com
  24. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback(意見截止 2026-10-19). U.S. FDA, 2026-08. fda.gov
  25. HHS announces new efforts to speed up, expand clinical trials with AI. STAT, 2026-09-30(付費牆/paywalled). statnews.com
  26. 生技論壇 10 月 7 日登場 衛福部長石崇良:積極建構 AI 醫療基建. 經濟日報, 2026-09-15. money.udn.com
  27. 食品藥物管理署智慧醫療器材專案辦公室成立. 衛生福利部, 2021-05-07. mohw.gov.tw
  28. 智慧醫療器材資訊暨媒合平台. 衛生福利部食品藥物管理署. aimd.fda.gov.tw
  29. 衛生福利部三大 AI 中心啟動記者會. 衛生福利部, 2024-10-07. mohw.gov.tw
  30. Everything That Happened in AI Today (Thursday, October 1 2026). The Neuron, 2026-10-01. theneuron.ai
Editor's note: This is a weekly review covering 28 September to 4 October 2026, Taipei time. A few items reach back to 22 September for comparison (Abridge's VA contract, the Heidi and Basalt rounds) and are dated individually in the text. Three STAT reports (Anthropic joining ARPA-H, the Euan Ashley opinion piece, HHS on clinical trials) sit behind a paywall and are written up only from the publicly visible headline, lede and summary; full text was not obtained, so Anthropic's specific role and contribution within ADVOCATE cannot be confirmed. Anthropic's IPO timing and the $2tn valuation come entirely from press reports citing unnamed sources and are not confirmed by the company. AKASA's 90 seconds, 6x annual growth, 1-in-10 inpatient discharge share and $180bn in client net patient revenue, along with Welldoc's 500m data points and Heidi's valuation, are all vendor-reported and not independently audited; AKASA's cited third-party validation names no validator and publishes no method or sample size, and is not peer-reviewed. The Nature Medicine emergency-department LLM study (19 August), the general-versus-specialised head-to-head and the April editorial were not published this week and are included as methodology for calibrating this week's claims, with their original dates stated in the text. McKinsey's 16–22% is a ceiling on technical feasibility rather than a deployment rate, and rests on 2024 claims data. Budget figures and timelines in the Taiwan section are quoted from the minister's remarks as reported by Economic Daily News on 15 September; the 7 October forum had not yet taken place at the time of writing. Not investment or medical advice.