◆ AI & Medical AI Daily
–
Thursday · Technology Breakthroughs

Reasoning is leaving the screen: Astra folds it into latent space, Gemini spreads it over cheap tokens, and a BCI paper prices the cost — 28.4% of high-confidence outputs are unfaithful

Three technical developments this week converge on one question: can humans still see a model reason? On September 1, TechCrunch reported that OpenAI's Astra is the first model to cross the company's own "critical cybersecurity threshold," scoring a perfect result on ExploitBench and autonomously discovering and exploiting two previously unknown zero-days in a modified evaluation (TechCrunch, 2026-09-01); in parallel, the claim circulating among researchers is that it uses "recurrent depth," keeping part of its thinking in hidden state rather than writing it out as chain-of-thought. A day later Google took the opposite road: Gemini 3.8 Flash and a security-focused Flash Cyber variant shipped together, cutting the price of explicit reasoning to $0.75 per million input tokens (Artificial Analysis, 2026-09-02). And on the clinical side, a medRxiv preprint posted August 27 converts invisible reasoning into a patient-level cost: when LLMs "correct" typing from brain–computer interface users, semantic drift reaches 60.3% at a 40% decoder error rate, and 28.4% of high-confidence outputs are unfaithful to the intended message (medRxiv, 2026-08-27).

01 — Top Stories

Seven items on two layers: the first four decide how models think, the last three decide what hardware stands in the operating room
Model architecture OpenAIAstra9/01

Astra becomes the first model to cross OpenAI's own "critical cybersecurity threshold" — a perfect ExploitBench score and two self-found zero-days

What

OpenAI says Astra can locate and exploit unknown vulnerabilities in computer systems without human guidance, scoring a perfect result on ExploitBench — an evaluation of LLM hacking ability — and surfacing two previously unknown zero-days in OpenAI's modified version of the test. The company says it plans to make Astra available "soon" while restricting its most advanced cybersecurity capabilities at first, paired with improved abuse detection, jailbreak prevention, risk-tiered account restrictions and chain-of-thought monitoring (TechCrunch, 2026-09-01). The Astra family was first announced on August 1 and updated on August 6 with solutions to ten long-unsolved mathematical problems, all formalised in Lean at roughly $2,000 of API tokens (The Decoder, 2026-08).

Why it matters

Hospitals are among the most poorly defended large organisations on the network, and a model that finds zero-days on its own is a tool for defenders and attackers alike. More pointed: OpenAI itself lists chain-of-thought monitoring among its safeguards — and if the recurrent-depth reports are accurate, there is less chain of thought left to monitor. TechCrunch also notes the safety claims currently lack independent third-party verification.

Discount this

"Recurrent depth" so far rests on community readings and secondary coverage; OpenAI has published no checkable technical report. The perfect ExploitBench score and the two zero-days are vendor-reported and not independently audited. Sebastian Raschka's September analysis pushes back on the idea that looping necessarily suppresses visible reasoning, arguing that reusing layers does not by itself hide a chain of thought (Raschka, 2026).

Model release Google DeepMindGemini 3.8 Flash9/02

Google ships Gemini 3.8 Flash and Flash Cyber on the same day: an intelligence index of 59 at the frontier, priced down to $0.75 per million tokens

What

Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index at high reasoning — three points above 3.7 Flash and level with GPT-5.6 Sol and Grok 4.6 — with a 1M-token context window, text/image/video/speech input and roughly 300 output tokens per second. Pricing is discounted to $0.75/$3.75 per million input/output tokens through the end of 2026, with a standard rate of $1.50/$7.50 and about $0.58 per task. It is Google's fourth Flash model in under four months (Artificial Analysis, 2026-09-02). The Flash Cyber variant, launched the same day for security researchers finding vulnerabilities and writing patches, beat Claude Opus 5 and GPT-5.6 Sol on 9 of 16 tests, scoring 73.7% on DeepSWE-1.1 and 86.2% on CyberGym C/C++ vulnerability detection (SiliconANGLE, 2026-09-02).

Why it matters

For health systems, what decides whether AI reaches the whole hospital is usually unit cost rather than peak capability: at the discounted rate, sweeping a 200,000-token inpatient chart costs roughly $0.15 in input tokens. Note also the governance split — Google turned the security capability into a public, purchasable variant while OpenAI restricted access to its own. Google's Tulsee Doshi and Raluca Ada Popa say the model "exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively" on complex tasks.

Discount this

The benchmark figures come from Google's own reporting and a third-party index, and none of them are medical tasks; no clinical metrics such as MedQA or image interpretation accompany this release. The discount expires at the end of 2026, so long-run budgeting should assume the $1.50/$7.50 standard rate.

Clinical evaluation medRxivBCI8/27

When an LLM "fixes" an ALS patient's sentence: 60.3% semantic drift at a 40% decoder error rate, and 28.4% of high-confidence outputs are unfaithful

What

A team drawn largely from radiology and informatics (Gorenshtein, Omar, Klang, Barash and colleagues) posted "Intent Drift in LLM-Assisted BCI Communication" on medRxiv, simulating P300-speller decoder corruption to test how faithfully 20 open-weight models repair text for brain–computer interface users. Drift climbed from 2.2% with no decoder error to 60.3% at high corruption; 28.4% of the models' own high-confidence outputs were unfaithful to the intended message; discrimination was moderate (AUROC 0.83) but calibration was poor (expected calibration error 0.32); the best correction policy still produced wrong messages in 18% of cases; clinically critical messages drifted slightly more, not less; and physician review agreed with automated scoring only moderately (kappa 0.41) (medRxiv, 2026-08-27).

Why it matters

This puts a number on model confabulation in the one setting where it is least tolerable: the user is a patient who has lost the ability to speak, and the output is taken as their own words. The authors' phrase is "fluent, highly confident lies" — the failure is not an uncorrected typo but a changed intent, asserted with conviction. An ECE of 0.32 means the confidence score is close to useless as a gate, which invalidates an entire class of deployment designs built on confidence thresholds.

Discount this

This is an in-silico benchmark: decoder errors were simulated rather than drawn from real BCI users, and it is a preprint that has not been peer reviewed. All 20 models are open-weight, with frontier closed models (Astra, Gemini 3.8, Claude Opus 5) excluded, so the figures cannot be extrapolated directly to the commercial models actually deployed in clinics.

AI agent ScreenAgentmeta-analysis8/30

Screening 200,000 papers for $855.91: agent–human agreement (kappa 0.75) now exceeds human–human agreement (0.64)

What

An LLM agent called ScreenAgent was built to automate the first-pass screening stage of medical meta-analyses. On the reported figures it processed 201,064 records for a total of $855.91 — about 0.43 cents each — with an internal sensitivity of 97.7% (43 of 44 eligible studies found) and a 99.4% workload reduction for human reviewers. Agent-to-human agreement reached a Cohen kappa of 0.75, above the 0.64 measured between human reviewers, with external validation sensitivities of 95.9% and 97.4%. The authors use a hybrid design in which a second LLM verifies results and human experts retain the final call (Yesil Science, 2026-08-30).

Why it matters

Systematic review is the cost bottleneck of evidence-based medicine; first-pass screening for a large meta-analysis typically occupies two researchers for months. If the 0.43-cents-per-record figure holds, the refresh cycle for clinical guidelines stops being limited by labour and starts being limited by how fast primary research appears. The contrast with the previous item is instructive: the same class of technology is badly calibrated when speaking on a patient's behalf, yet robust on a task with an explicit recall target and a second pass to catch its errors. Task structure, not raw model strength, decides whether an agent is usable.

Discount this

We could only reach secondary coverage for this item and could not verify the numbers against the medRxiv original, so every figure above rests on that report. The system missed one eligible study in testing — and in a meta-analysis, the missed study can be the one that changes the conclusion. A 97.7% sensitivity sounds high, but evidence synthesis conventionally demands a stricter bar.

Imaging hardware GE HealthCarePhotonova Spectra8/31

GE HealthCare's photon-counting CT clears CE marking: the Deep Silicon detector measures individual photon energies directly, skipping the scintillation step

What

GE HealthCare's Photonova Spectra received CE marking on August 31. Where conventional CT detectors first convert X-rays into visible light, the Deep Silicon detector measures individual X-ray photons and their energies directly, so every scan yields high-definition images and spectral data at once, across neurology, oncology, musculoskeletal, thoracic and cardiac imaging. The system already held FDA 510(k) clearance and Japanese approval as of March 2026. Chad Rowland, executive director of global CT at GE HealthCare, said: "At its core is our Deep Silicon detector technology, which enables Photonova Spectra to capture remarkably detailed, spectral information in every scan." (MobiHealthNews, 2026-08-31)

Why it matters

In a week dominated by software, this is one of the few breakthroughs at the physics layer — and it determines what the models above it get to see. Spectral CT carries material composition in every voxel rather than a single HU value, which effectively swaps imaging AI's input from a greyscale picture to a multi-channel measurement. Imaging foundation models trained on conventional CT cannot simply ingest that data, and the cost of retraining and revalidation lands on the hospitals that buy photon-counting scanners.

Discount this

The coverage provides no figures for resolution, dose reduction or clinical outcomes, and no head-to-head comparison against existing photon-counting platforms such as Siemens' Naeotom. CE marking is a licence to sell, not evidence of clinical benefit.

Surgical tech StrykerApple Vision Pro9/01

The first FDA De Novo-authorised intraoperative Apple Vision Pro application: Stryker's SportSuite Vision completes a first hip arthroscopy at Duke

What

Stryker announced the first clinical use of SportSuite Vision on September 1. It is a video see-through augmented reality head-mounted display that folds arthroscopic video, the HipCheck software, HipMap femoroacetabular impingement analysis and CT imaging into a customisable spatial computing environment, overlaying digital clinical information inside the surgeon's direct field of view. The FDA granted De Novo authorisation on July 17, 2026, the first approval for intraoperative use with Apple Vision Pro; the inaugural case was performed by Duke Health orthopaedic surgeon Chad Mather III, with the indication limited to hip arthroscopy for FAI and labral repair (Stryker, 2026-09-01; STAT, 2026-09-01).

Why it matters

De Novo rather than 510(k) means the FDA found no comparable predicate, so this creates a new classification and special controls for intraoperative head-mounted displays — later entrants can follow via 510(k). Dr Mather's account points to ergonomics rather than immersion as the real value: "spatial computing enabled me to customize the placement of key clinical information to fit my workflow and access it within the sterile field." This is the first time consumer hardware has entered the sterile field as the primary display.

Discount this

Stryker's announcement mentions no AI component — this is a display and integration system, not automated interpretation. The surgeon on the first case is also a Stryker consultant, and with a single case there are no comparative data on operative time, complications or learning curve. Most of the STAT piece sits behind a paywall, so this item rests chiefly on Stryker's own press release.

Surgical robotics MedtronicCornerstone Robotics9/01

Medtronic pays $700M for Sentire's ex-US distribution rights: running two robot platforms at once is a bet on the connected surgical ecosystem, not a single machine

What

Medtronic announced on September 1 an investment of roughly $700 million for distribution rights to Hong Kong-based Cornerstone Robotics' Sentire surgical system in select approved markets outside the United States. Sentire was developed entirely in-house, features dual-console capability and an immersive console design, received CE marking in May 2026 for minimally invasive general, gynaecologic, thoracic and urologic procedures, and holds market approval in China and Singapore. Medtronic's own Hugo RAS operates in more than 35 countries across six continents, with global procedures projected to exceed 50,000 by fiscal year end, and a digital ecosystem that includes real-time AI and the Touch Surgery platform (Medtronic, 2026-09-01).

Why it matters

Value in surgical robotics is migrating from the arm to the data it generates: every installed system is a continuous source of operative video and instrument kinematics, which is the one training corpus for surgical AI that cannot easily be substituted. By running Hugo and Sentire together, Medtronic trades distribution for data scale, in contrast to Intuitive Surgical's single-platform strategy — making it the second player outside the da Vinci franchise with a plausible path to pooling data across machine types. Matt Anderson, senior vice president of Medtronic's surgical division, said "surgeons, health systems and markets have varying needs and preferences," calling for flexibility across clinical settings.

Discount this

The release does not say whether data generated by Sentire feeds Medtronic's digital ecosystem, nor how the $700 million splits between equity and distribution rights, nor Sentire's installed base or procedure volume. The 50,000 figure is a Hugo projection, not a realised number.

02 — Product Analysis

Two models shipped in the same week embody opposite answers to where reasoning should live — and that choice decides whether medicine can use them

OpenAI Astra

A frontier model that folds reasoning into latent space · OpenAI (US)

Function and position. Astra is positioned as a model family for long-horizon, complex problems, coordinating multiple agents over hours or days. The ten previously unsolved mathematical results published on August 6 were all formalised in Lean at roughly $2,000 in API tokens, and mathematician Thomas Bloom called them "more significant than the counterexample to the unit distance conjecture published in May" (The Decoder). It is not yet generally available, and whether it ships as GPT-6 or a GPT-5 variant remains undecided.

  • Strength : the capability ceiling genuinely moved — a perfect ExploitBench score, two self-found zero-days, ten long-unsolved mathematical problems (TechCrunch). Reusing one layer stack via recurrent depth, rather than stacking new layers, buys deeper effective computation without adding parameters.
  • Concern : clinical deployment needs an auditable reasoning trace — Taiwan's MOHW generative AI guidance and the FDA's transparency expectations both assume a model can account for itself. If part of the reasoning exists only in hidden state, keeping a record becomes technically hard. And there is no official technical report to check any of this against: even whether Astra works this way remains community inference.
  • Concern : security researcher Yona Shavit questions whether Astra's refusal to break containment reflects genuine alignment or deliberate deception; OpenAI's safety claims lack independent third-party verification, and the status of any government collaboration is unclear (TechCrunch).

Gemini 3.8 Flash / Flash Cyber

Volume models that make explicit reasoning cheap · Google DeepMind (US)

Function and position. The Flash line pursues affordable frontier performance: an intelligence index of 59, a 1M-token context, and three reasoning levels (high/medium/low) at 2.5 minutes per task on high and 0.8 on low, letting the deployer set the dial between cost and quality (Artificial Analysis, 2026-09-02). Flash Cyber packages the security capability as a separate, openly sold variant.

  • Strength : adjustable reasoning depth is an underrated feature for medicine — low for triage, high for hard cases, cost tiering through one API. At the discounted $0.75/$3.75, hospital-wide deployment enters budget range for the first time (Artificial Analysis).
  • Strength : Flash Cyber scores 86.2% on CyberGym C/C++ vulnerability detection and 73.7% on DeepSWE-1.1, beating Claude Opus 5 and GPT-5.6 Sol on 9 of 16 tests (SiliconANGLE). Medical device software is heavily legacy C/C++, which makes this immediately applicable.
  • Concern : the release carries no medical benchmarks at all. Four Flash models in four months is a burden rather than a benefit for hospitals — every version change demands revalidation, and the clinical safety validation and compatibility testing Taiwan's MOHW guidance requires cannot run on a three-week release cadence. Procurement should also price in the doubling when the discount expires at year end.

Where the contrast lies. Astra buys ceiling by hiding computation in latent space; Gemini buys affordability and observability by spreading computation over cheap tokens. What medicine can currently both afford and audit is the latter — while the direction the former represents makes the regulatory premise that a model must be able to explain itself progressively harder to satisfy in engineering terms.

03 — Companies & Competition

Who stands where, on what, against whom
Company Recent state & numbers Position & moat
OpenAI
Definer of the capability frontier
Astra announced 8/01, with ten unsolved mathematical results added 8/06 (formalised in Lean, ~$2,000 of API tokens); on 9/01 confirmed as the first model past the critical cybersecurity threshold with a perfect ExploitBench score (TechCrunch, 2026-09-01). The moat is capability lead and research brand. The weakness is transparency: no technical report, no third-party audit, and the stronger the capability the more access must be restricted — which squeezes out markets like medicine that require full inspection.
Google DeepMind
The one applying price and scale pressure
Shipped Gemini 3.8 Flash (index 59, 1M context, discounted $0.75/$3.75) and Flash Cyber (86.2% on CyberGym) on 9/02 — its fourth Flash model in four months (Artificial Analysis, 2026-09-02). On the clinical side it also holds the open-weight MedGemma 1.5 line (2026-01-13, 4B multimodal, 69.1% MedQA, 89.6% EHR understanding) (Google HAI-DEF). The only player holding both a frontier general model and open-weight medical models, letting hospitals run MedGemma on sensitive data in-house and Flash for volume work. The weakness is a release cadence badly out of phase with hospital validation cycles.
GE HealthCare
Imaging hardware plus AI software
Photonova Spectra photon-counting CT received CE marking on 8/31, having held FDA 510(k) clearance and Japanese approval since March 2026 (MobiHealthNews, 2026-08-31). On the radiotherapy side, MIM Contour ProtégéAI+ 2.0 gained FDA 510(k) clearance with a PCCP on 2026-06-08 (Applied Radiation Oncology). The moat is owning where data is generated: the detector defines the input format for imaging AI, and a PCCP lets models update without a fresh 510(k) each time. Rivals are Siemens Healthineers' Naeotom photon-counting platform and Philips. The weakness is a hardware replacement cycle measured in years against software rivals iterating weekly.
Stryker
Bringing consumer hardware into the sterile field
SportSuite Vision received FDA De Novo authorisation on 2026-07-17 and completed its first hip arthroscopy at Duke Health on 9/01, integrating HipCheck, HipMap FAI analysis and CT (Stryker, 2026-09-01). The moat is the new classification its De Novo creates plus first-mover status, layered on an existing orthopaedic implant and instrument channel. The weakness is tying the core experience to a single consumer electronics platform — and the absence of any AI component: this is display-layer innovation, not interpretation.
Medtronic
Trading distribution for surgical data scale
Announced roughly $700 million on 9/01 for ex-US distribution rights to Cornerstone Robotics' Sentire; its own Hugo RAS is in more than 35 countries, with global procedures projected to exceed 50,000 this fiscal year (Medtronic, 2026-09-01). A two-platform strategy covers both premium and price-sensitive markets, with the Touch Surgery ecosystem as the data confluence. The rival is Intuitive Surgical's installed da Vinci base. The weakness: whether the two platforms' data interoperate is unstated — and if they do not, the distribution advantage never converts into a training-data advantage.
Cornerstone Robotics
Hong Kong's in-house surgical robot
Sentire was developed entirely in-house with a dual-console design, received CE marking in May 2026 and holds approvals in China and Singapore; the company runs three global R&D hubs and six business centres (Medtronic, 2026-09-01). The moat is regulatory progress and cost structure in Asian markets, plus global distribution acquired through the Medtronic tie-up. The weakness is handing ex-US distribution to a company that simultaneously sells a competing platform of its own, which caps long-run bargaining power.
Elucid
AI plaque characterisation in cardiovascular imaging
Announced a $55M Series D on 9/02, taking total funding to about $185M, with investors including an unnamed large public medtech company, IAG Capital Partners and Elevage Medical Technologies; Plaque-IQ is FDA cleared while its FFR-CT version remains under review (MobiHealthNews, 2026-09-02). Its link to today's through-line: plaque composition analysis is one of the applications spectral photon-counting data most directly enables. The moat is FDA clearance and embedding in existing CT angiography workflows, against HeartFlow and Cleerly. The absence of disclosed outcome data is a conspicuous gap.

Today's competitive structure is stacked, not head-to-head. The model layer (OpenAI, Google) is fighting over what "reasoning" should look like; the hardware layer (GE, Stryker, Medtronic, Cornerstone) is fighting over who owns the point where data is generated. What is genuinely scarce is neither parameters nor robotic arms but an auditable reasoning record on one side and a real clinical data stream fit for training on the other — and across today's seven items, no single company holds both.

04 — Taiwan Angle

The six risks Taiwan's MOHW named in May were hit one by one this week

(1) Hallucination and prompt injection stopped being abstract clauses. On May 29, 2026, Taiwan's Ministry of Health and Welfare issued its guidance on generative AI in medical institutions, naming six risks — two of which, "AI hallucination producing false information" and "prompt injection attacks threatening cybersecurity," acquired concrete numbers this week: the BCI paper measured 28.4% of high-confidence outputs as unfaithful, and Astra demonstrated a model finding and exploiting zero-days on its own. The guidance is advisory rather than binding; its nine requirements span pre-deployment assignment of responsibility and data reliability assessment, integration-phase compatibility testing against EHR and imaging systems plus clinical safety validation, and post-deployment retention of final clinical authority by medical professionals, output quality monitoring, bias management and informing patients when AI participates in their care (GlobalBio Monthly).

(2) The three MOHW AI centres map neatly onto what each of today's stories is missing. Taiwan operates a Center for Responsible AI in Healthcare, a Center for External AI Validation in Healthcare (linking cross-hospital data via FHIR and running a federated learning platform) and a Center for Clinical AI Impact Evaluation (health economics analysis and National Health Insurance pricing) (Taiwan's three smart healthcare centres). Line today's cases up against them: an agent like ScreenAgent lacks the external validation the second centre exists to perform; Gemini 3.8 Flash lacks the cost-effectiveness analysis the third would run; and Astra lacks exactly the auditability — regular evaluation, a completed data science cycle — that the first is built around. Three gaps, three centres, and not one of their processes was designed for a model version that changes every three weeks.

(3) Photon-counting CT is the procurement decision Taiwanese hospitals will face soonest. With FDA, Japanese and CE approvals all in hand, Photonova Spectra reaching Taiwan is a matter of timing. But there is an easily overlooked knock-on cost: spectral data is not formatted like conventional CT, so the in-house AI models a hospital has already trained and validated on conventional CT images cannot simply be pointed at the new scanner's output. Buying a photon-counting CT therefore also means sending a batch of already-validated in-house models back for revalidation — a cost that belongs in the procurement assessment, not discovered once the queue at the external validation centre has formed.

05 — Further Reading

Chosen to outlast this week's news cycle: one primary paper, two large-scale clinical foundation models, and one piece of technical clarification
  1. Intent Drift in LLM-Assisted BCI Communication: An In-Silico Benchmark Under Simulated Decoder Corruption — medRxiv (2026-08-27)

    The one item today worth reading in full: it measures "confidence scores cannot serve as a gate" to a precision you could write into a deployment spec, and the method transfers directly to faithfulness evaluation for ambient scribes.

  2. Apollo: A multimodal and temporal foundation model for virtual patient representations at healthcare system scale — arXiv (2026-04-20)

    Faisal Mahmood's group trained on 25 billion records from 7.2 million patients across 28 modalities and evaluated on 1.4 million patients over 322 prediction tasks. The most complete reference point available for what a record-scale foundation model actually looks like.

  3. Learning from routine health system data builds better neuroimaging AI models — Nature Medicine (2026-07-31)

    A visual foundation model trained on 5.24 million routine clinical CT and MRI series outperformed models trained on public internet and medical datasets — an answer running against the common assumption that messy in-house data is a liability.

  4. OpenAI Astra and Looped Transformers — Sebastian Raschka (2026-09)

    Amid a wave of "Astra is hiding its reasoning" takes, this explains what layer looping actually does and argues the opposite. Read it before accepting today's through-line at face value.

  5. MedGemma 1.5 model card — Google Health AI Developer Foundations (2026-01-13)

    If today's through-line makes you uneasy about frontier model auditability, this is the alternative you can deploy in-house and check benchmark by benchmark; the jump in EHR understanding from 67.6% to 89.6% is worth reading alongside the terms of use.

06 — References

References
  1. OpenAI's Astra model is on the way — and very good at breaking into computer systems. TechCrunch, 2026-09-01. techcrunch.com
  2. OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions. The Decoder, 2026-08. the-decoder.com
  3. OpenAI Astra and Looped Transformers. Sebastian Raschka, 2026-09. sebastianraschka.com
  4. Google has released Gemini 3.8 Flash, its fourth Flash model in under four months. Artificial Analysis, 2026-09-02. artificialanalysis.ai
  5. Google launches two Gemini 3.8 models with cutting-edge reasoning capabilities. SiliconANGLE, 2026-09-02. siliconangle.com
  6. Gorenshtein A, Omar M, Jia E, Adiniaev Y, Daniel O, Kruskal J, Ahmed M, Brook O, Klang E, Barash Y. Intent Drift in LLM-Assisted BCI Communication: An In-Silico Benchmark Under Simulated Decoder Corruption. medRxiv, 2026-08-27. medrxiv.org
  7. AI Screens 200,000 Medical Papers for Pennies (ScreenAgent). Yesil Science, 2026-08-30. yesilscience.com
  8. The Health AI Brief — Week of August 31, 2026. Yesil Science, 2026-08-31. yesilscience.com
  9. GE HealthCare secures CE mark for its Photonova Spectra technology. MobiHealthNews, 2026-08-31. mobihealthnews.com
  10. FDA Clears GE HealthCare's MIM Contour ProtégéAI+ 2.0 for Advanced Cancer Treatment Planning. Applied Radiation Oncology, 2026-06-08. appliedradiationoncology.com
  11. Stryker introduces first of its kind FDA-authorized surgical application for Apple Vision Pro with hip arthroscopy case. Stryker Newsroom, 2026-09-01. stryker.com
  12. FDA clears Stryker's 'surgical cockpit' that uses Apple's VR headset, Vision Pro. STAT News, 2026-09-01. statnews.com
  13. Medtronic Announces Strategic Partnership with Cornerstone Robotics to Further Expand Global Access to Robotic-Assisted Surgery. Medtronic Newsroom, 2026-09-01. news.medtronic.com
  14. Medtronic invests $700M in Cornerstone Robotics. MobiHealthNews, 2026-09-01. mobihealthnews.com
  15. Elucid raises $55M for AI cardiovascular imaging software. MobiHealthNews, 2026-09-02. mobihealthnews.com
  16. MedGemma 1.5 model card. Google Health AI Developer Foundations, 2026-01-13. developers.google.com
  17. Apollo: A multimodal and temporal foundation model for virtual patient representations at healthcare system scale. arXiv:2604.18570, 2026-04-20. arxiv.org
  18. Learning from routine health system data builds better neuroimaging AI models. Nature Medicine 32(8), 2026-07-31. nature.com
  19. 衛福部首發醫療 GenAI 指引:六大風險、九項要點建立治理框架. 環球生技月刊, 2026-05-29. news.gbimonthly.com
  20. 臺灣智慧醫療三大中心. 衛生福利部. aicenter.mohw.gov.tw
  21. AI model release log (Gemini 3.8 Flash, Claude Fable 5.1, GLM-5.3-Flash, Qwen3.8 Flash, Granite 4.2). LLM-Stats, 2026-09. llm-stats.com
Editor's note: (1) Most of the STAT News piece on Stryker sits behind a paywall; that item rests chiefly on Stryker's own release, and paywalled content was not used. (2) For ScreenAgent we obtained only secondary coverage from Yesil Science and could not verify against the medRxiv original — every figure in that item (201,064 records, $855.91, 97.7% sensitivity, kappa 0.75) rests on that report, and readers should check the primary paper before citing. (3) Astra's "recurrent depth" architecture comes from community readings and secondary coverage; OpenAI has published no checkable technical report, and the perfect ExploitBench score and two zero-days are vendor-reported and unaudited. (4) The Gemini 3.8 Flash benchmarks mix Google's own claims with the third-party Artificial Analysis index, and include no medical task evaluation. (5) The MOHW guidance cited in section 4 was issued on 2026-05-29 and is not this week's news; it is used as a background framework. (6) GE HealthCare's MIM Contour ProtégéAI+ 2.0 was cleared on 2026-06-08 and only resurfaced this week in secondary coverage, so it is not carried as a top story. (7) Today's rotation is technology, so Elucid's financing appears only in the company table for its technical link to photon-counting CT rather than as a standalone item.