Reasoning is leaving the screen: Astra folds it into latent space, Gemini spreads it over cheap tokens, and a BCI paper prices the cost — 28.4% of high-confidence outputs are unfaithful
Three technical developments this week converge on one question: can humans still see a model reason? On September 1, TechCrunch reported that OpenAI's Astra is the first model to cross the company's own "critical cybersecurity threshold," scoring a perfect result on ExploitBench and autonomously discovering and exploiting two previously unknown zero-days in a modified evaluation (TechCrunch, 2026-09-01); in parallel, the claim circulating among researchers is that it uses "recurrent depth," keeping part of its thinking in hidden state rather than writing it out as chain-of-thought. A day later Google took the opposite road: Gemini 3.8 Flash and a security-focused Flash Cyber variant shipped together, cutting the price of explicit reasoning to $0.75 per million input tokens (Artificial Analysis, 2026-09-02). And on the clinical side, a medRxiv preprint posted August 27 converts invisible reasoning into a patient-level cost: when LLMs "correct" typing from brain–computer interface users, semantic drift reaches 60.3% at a 40% decoder error rate, and 28.4% of high-confidence outputs are unfaithful to the intended message (medRxiv, 2026-08-27).
01 — Top Stories
Astra becomes the first model to cross OpenAI's own "critical cybersecurity threshold" — a perfect ExploitBench score and two self-found zero-days
OpenAI says Astra can locate and exploit unknown vulnerabilities in computer systems without human guidance, scoring a perfect result on ExploitBench — an evaluation of LLM hacking ability — and surfacing two previously unknown zero-days in OpenAI's modified version of the test. The company says it plans to make Astra available "soon" while restricting its most advanced cybersecurity capabilities at first, paired with improved abuse detection, jailbreak prevention, risk-tiered account restrictions and chain-of-thought monitoring (TechCrunch, 2026-09-01). The Astra family was first announced on August 1 and updated on August 6 with solutions to ten long-unsolved mathematical problems, all formalised in Lean at roughly $2,000 of API tokens (The Decoder, 2026-08).
Hospitals are among the most poorly defended large organisations on the network, and a model that finds zero-days on its own is a tool for defenders and attackers alike. More pointed: OpenAI itself lists chain-of-thought monitoring among its safeguards — and if the recurrent-depth reports are accurate, there is less chain of thought left to monitor. TechCrunch also notes the safety claims currently lack independent third-party verification.
"Recurrent depth" so far rests on community readings and secondary coverage; OpenAI has published no checkable technical report. The perfect ExploitBench score and the two zero-days are vendor-reported and not independently audited. Sebastian Raschka's September analysis pushes back on the idea that looping necessarily suppresses visible reasoning, arguing that reusing layers does not by itself hide a chain of thought (Raschka, 2026).
Google ships Gemini 3.8 Flash and Flash Cyber on the same day: an intelligence index of 59 at the frontier, priced down to $0.75 per million tokens
Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index at high reasoning — three points above 3.7 Flash and level with GPT-5.6 Sol and Grok 4.6 — with a 1M-token context window, text/image/video/speech input and roughly 300 output tokens per second. Pricing is discounted to $0.75/$3.75 per million input/output tokens through the end of 2026, with a standard rate of $1.50/$7.50 and about $0.58 per task. It is Google's fourth Flash model in under four months (Artificial Analysis, 2026-09-02). The Flash Cyber variant, launched the same day for security researchers finding vulnerabilities and writing patches, beat Claude Opus 5 and GPT-5.6 Sol on 9 of 16 tests, scoring 73.7% on DeepSWE-1.1 and 86.2% on CyberGym C/C++ vulnerability detection (SiliconANGLE, 2026-09-02).
For health systems, what decides whether AI reaches the whole hospital is usually unit cost rather than peak capability: at the discounted rate, sweeping a 200,000-token inpatient chart costs roughly $0.15 in input tokens. Note also the governance split — Google turned the security capability into a public, purchasable variant while OpenAI restricted access to its own. Google's Tulsee Doshi and Raluca Ada Popa say the model "exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively" on complex tasks.
The benchmark figures come from Google's own reporting and a third-party index, and none of them are medical tasks; no clinical metrics such as MedQA or image interpretation accompany this release. The discount expires at the end of 2026, so long-run budgeting should assume the $1.50/$7.50 standard rate.
When an LLM "fixes" an ALS patient's sentence: 60.3% semantic drift at a 40% decoder error rate, and 28.4% of high-confidence outputs are unfaithful
A team drawn largely from radiology and informatics (Gorenshtein, Omar, Klang, Barash and colleagues) posted "Intent Drift in LLM-Assisted BCI Communication" on medRxiv, simulating P300-speller decoder corruption to test how faithfully 20 open-weight models repair text for brain–computer interface users. Drift climbed from 2.2% with no decoder error to 60.3% at high corruption; 28.4% of the models' own high-confidence outputs were unfaithful to the intended message; discrimination was moderate (AUROC 0.83) but calibration was poor (expected calibration error 0.32); the best correction policy still produced wrong messages in 18% of cases; clinically critical messages drifted slightly more, not less; and physician review agreed with automated scoring only moderately (kappa 0.41) (medRxiv, 2026-08-27).
This puts a number on model confabulation in the one setting where it is least tolerable: the user is a patient who has lost the ability to speak, and the output is taken as their own words. The authors' phrase is "fluent, highly confident lies" — the failure is not an uncorrected typo but a changed intent, asserted with conviction. An ECE of 0.32 means the confidence score is close to useless as a gate, which invalidates an entire class of deployment designs built on confidence thresholds.
This is an in-silico benchmark: decoder errors were simulated rather than drawn from real BCI users, and it is a preprint that has not been peer reviewed. All 20 models are open-weight, with frontier closed models (Astra, Gemini 3.8, Claude Opus 5) excluded, so the figures cannot be extrapolated directly to the commercial models actually deployed in clinics.
Screening 200,000 papers for $855.91: agent–human agreement (kappa 0.75) now exceeds human–human agreement (0.64)
An LLM agent called ScreenAgent was built to automate the first-pass screening stage of medical meta-analyses. On the reported figures it processed 201,064 records for a total of $855.91 — about 0.43 cents each — with an internal sensitivity of 97.7% (43 of 44 eligible studies found) and a 99.4% workload reduction for human reviewers. Agent-to-human agreement reached a Cohen kappa of 0.75, above the 0.64 measured between human reviewers, with external validation sensitivities of 95.9% and 97.4%. The authors use a hybrid design in which a second LLM verifies results and human experts retain the final call (Yesil Science, 2026-08-30).
Systematic review is the cost bottleneck of evidence-based medicine; first-pass screening for a large meta-analysis typically occupies two researchers for months. If the 0.43-cents-per-record figure holds, the refresh cycle for clinical guidelines stops being limited by labour and starts being limited by how fast primary research appears. The contrast with the previous item is instructive: the same class of technology is badly calibrated when speaking on a patient's behalf, yet robust on a task with an explicit recall target and a second pass to catch its errors. Task structure, not raw model strength, decides whether an agent is usable.
We could only reach secondary coverage for this item and could not verify the numbers against the medRxiv original, so every figure above rests on that report. The system missed one eligible study in testing — and in a meta-analysis, the missed study can be the one that changes the conclusion. A 97.7% sensitivity sounds high, but evidence synthesis conventionally demands a stricter bar.
GE HealthCare's photon-counting CT clears CE marking: the Deep Silicon detector measures individual photon energies directly, skipping the scintillation step
GE HealthCare's Photonova Spectra received CE marking on August 31. Where conventional CT detectors first convert X-rays into visible light, the Deep Silicon detector measures individual X-ray photons and their energies directly, so every scan yields high-definition images and spectral data at once, across neurology, oncology, musculoskeletal, thoracic and cardiac imaging. The system already held FDA 510(k) clearance and Japanese approval as of March 2026. Chad Rowland, executive director of global CT at GE HealthCare, said: "At its core is our Deep Silicon detector technology, which enables Photonova Spectra to capture remarkably detailed, spectral information in every scan." (MobiHealthNews, 2026-08-31)
In a week dominated by software, this is one of the few breakthroughs at the physics layer — and it determines what the models above it get to see. Spectral CT carries material composition in every voxel rather than a single HU value, which effectively swaps imaging AI's input from a greyscale picture to a multi-channel measurement. Imaging foundation models trained on conventional CT cannot simply ingest that data, and the cost of retraining and revalidation lands on the hospitals that buy photon-counting scanners.
The coverage provides no figures for resolution, dose reduction or clinical outcomes, and no head-to-head comparison against existing photon-counting platforms such as Siemens' Naeotom. CE marking is a licence to sell, not evidence of clinical benefit.
The first FDA De Novo-authorised intraoperative Apple Vision Pro application: Stryker's SportSuite Vision completes a first hip arthroscopy at Duke
Stryker announced the first clinical use of SportSuite Vision on September 1. It is a video see-through augmented reality head-mounted display that folds arthroscopic video, the HipCheck software, HipMap femoroacetabular impingement analysis and CT imaging into a customisable spatial computing environment, overlaying digital clinical information inside the surgeon's direct field of view. The FDA granted De Novo authorisation on July 17, 2026, the first approval for intraoperative use with Apple Vision Pro; the inaugural case was performed by Duke Health orthopaedic surgeon Chad Mather III, with the indication limited to hip arthroscopy for FAI and labral repair (Stryker, 2026-09-01; STAT, 2026-09-01).
De Novo rather than 510(k) means the FDA found no comparable predicate, so this creates a new classification and special controls for intraoperative head-mounted displays — later entrants can follow via 510(k). Dr Mather's account points to ergonomics rather than immersion as the real value: "spatial computing enabled me to customize the placement of key clinical information to fit my workflow and access it within the sterile field." This is the first time consumer hardware has entered the sterile field as the primary display.
Stryker's announcement mentions no AI component — this is a display and integration system, not automated interpretation. The surgeon on the first case is also a Stryker consultant, and with a single case there are no comparative data on operative time, complications or learning curve. Most of the STAT piece sits behind a paywall, so this item rests chiefly on Stryker's own press release.
Medtronic pays $700M for Sentire's ex-US distribution rights: running two robot platforms at once is a bet on the connected surgical ecosystem, not a single machine
Medtronic announced on September 1 an investment of roughly $700 million for distribution rights to Hong Kong-based Cornerstone Robotics' Sentire surgical system in select approved markets outside the United States. Sentire was developed entirely in-house, features dual-console capability and an immersive console design, received CE marking in May 2026 for minimally invasive general, gynaecologic, thoracic and urologic procedures, and holds market approval in China and Singapore. Medtronic's own Hugo RAS operates in more than 35 countries across six continents, with global procedures projected to exceed 50,000 by fiscal year end, and a digital ecosystem that includes real-time AI and the Touch Surgery platform (Medtronic, 2026-09-01).
Value in surgical robotics is migrating from the arm to the data it generates: every installed system is a continuous source of operative video and instrument kinematics, which is the one training corpus for surgical AI that cannot easily be substituted. By running Hugo and Sentire together, Medtronic trades distribution for data scale, in contrast to Intuitive Surgical's single-platform strategy — making it the second player outside the da Vinci franchise with a plausible path to pooling data across machine types. Matt Anderson, senior vice president of Medtronic's surgical division, said "surgeons, health systems and markets have varying needs and preferences," calling for flexibility across clinical settings.
The release does not say whether data generated by Sentire feeds Medtronic's digital ecosystem, nor how the $700 million splits between equity and distribution rights, nor Sentire's installed base or procedure volume. The 50,000 figure is a Hugo projection, not a realised number.
02 — Product Analysis
OpenAI Astra
A frontier model that folds reasoning into latent space · OpenAI (US)
Function and position. Astra is positioned as a model family for long-horizon, complex problems, coordinating multiple agents over hours or days. The ten previously unsolved mathematical results published on August 6 were all formalised in Lean at roughly $2,000 in API tokens, and mathematician Thomas Bloom called them "more significant than the counterexample to the unit distance conjecture published in May" (The Decoder). It is not yet generally available, and whether it ships as GPT-6 or a GPT-5 variant remains undecided.
- Strength : the capability ceiling genuinely moved — a perfect ExploitBench score, two self-found zero-days, ten long-unsolved mathematical problems (TechCrunch). Reusing one layer stack via recurrent depth, rather than stacking new layers, buys deeper effective computation without adding parameters.
- Concern : clinical deployment needs an auditable reasoning trace — Taiwan's MOHW generative AI guidance and the FDA's transparency expectations both assume a model can account for itself. If part of the reasoning exists only in hidden state, keeping a record becomes technically hard. And there is no official technical report to check any of this against: even whether Astra works this way remains community inference.
- Concern : security researcher Yona Shavit questions whether Astra's refusal to break containment reflects genuine alignment or deliberate deception; OpenAI's safety claims lack independent third-party verification, and the status of any government collaboration is unclear (TechCrunch).
Gemini 3.8 Flash / Flash Cyber
Volume models that make explicit reasoning cheap · Google DeepMind (US)
Function and position. The Flash line pursues affordable frontier performance: an intelligence index of 59, a 1M-token context, and three reasoning levels (high/medium/low) at 2.5 minutes per task on high and 0.8 on low, letting the deployer set the dial between cost and quality (Artificial Analysis, 2026-09-02). Flash Cyber packages the security capability as a separate, openly sold variant.
- Strength : adjustable reasoning depth is an underrated feature for medicine — low for triage, high for hard cases, cost tiering through one API. At the discounted $0.75/$3.75, hospital-wide deployment enters budget range for the first time (Artificial Analysis).
- Strength : Flash Cyber scores 86.2% on CyberGym C/C++ vulnerability detection and 73.7% on DeepSWE-1.1, beating Claude Opus 5 and GPT-5.6 Sol on 9 of 16 tests (SiliconANGLE). Medical device software is heavily legacy C/C++, which makes this immediately applicable.
- Concern : the release carries no medical benchmarks at all. Four Flash models in four months is a burden rather than a benefit for hospitals — every version change demands revalidation, and the clinical safety validation and compatibility testing Taiwan's MOHW guidance requires cannot run on a three-week release cadence. Procurement should also price in the doubling when the discount expires at year end.
Where the contrast lies. Astra buys ceiling by hiding computation in latent space; Gemini buys affordability and observability by spreading computation over cheap tokens. What medicine can currently both afford and audit is the latter — while the direction the former represents makes the regulatory premise that a model must be able to explain itself progressively harder to satisfy in engineering terms.
03 — Companies & Competition
| Company | Recent state & numbers | Position & moat |
|---|---|---|
| OpenAI Definer of the capability frontier |
Astra announced 8/01, with ten unsolved mathematical results added 8/06 (formalised in Lean, ~$2,000 of API tokens); on 9/01 confirmed as the first model past the critical cybersecurity threshold with a perfect ExploitBench score (TechCrunch, 2026-09-01). | The moat is capability lead and research brand. The weakness is transparency: no technical report, no third-party audit, and the stronger the capability the more access must be restricted — which squeezes out markets like medicine that require full inspection. |
| Google DeepMind The one applying price and scale pressure |
Shipped Gemini 3.8 Flash (index 59, 1M context, discounted $0.75/$3.75) and Flash Cyber (86.2% on CyberGym) on 9/02 — its fourth Flash model in four months (Artificial Analysis, 2026-09-02). On the clinical side it also holds the open-weight MedGemma 1.5 line (2026-01-13, 4B multimodal, 69.1% MedQA, 89.6% EHR understanding) (Google HAI-DEF). | The only player holding both a frontier general model and open-weight medical models, letting hospitals run MedGemma on sensitive data in-house and Flash for volume work. The weakness is a release cadence badly out of phase with hospital validation cycles. |
| GE HealthCare Imaging hardware plus AI software |
Photonova Spectra photon-counting CT received CE marking on 8/31, having held FDA 510(k) clearance and Japanese approval since March 2026 (MobiHealthNews, 2026-08-31). On the radiotherapy side, MIM Contour ProtégéAI+ 2.0 gained FDA 510(k) clearance with a PCCP on 2026-06-08 (Applied Radiation Oncology). | The moat is owning where data is generated: the detector defines the input format for imaging AI, and a PCCP lets models update without a fresh 510(k) each time. Rivals are Siemens Healthineers' Naeotom photon-counting platform and Philips. The weakness is a hardware replacement cycle measured in years against software rivals iterating weekly. |
| Stryker Bringing consumer hardware into the sterile field |
SportSuite Vision received FDA De Novo authorisation on 2026-07-17 and completed its first hip arthroscopy at Duke Health on 9/01, integrating HipCheck, HipMap FAI analysis and CT (Stryker, 2026-09-01). | The moat is the new classification its De Novo creates plus first-mover status, layered on an existing orthopaedic implant and instrument channel. The weakness is tying the core experience to a single consumer electronics platform — and the absence of any AI component: this is display-layer innovation, not interpretation. |
| Medtronic Trading distribution for surgical data scale |
Announced roughly $700 million on 9/01 for ex-US distribution rights to Cornerstone Robotics' Sentire; its own Hugo RAS is in more than 35 countries, with global procedures projected to exceed 50,000 this fiscal year (Medtronic, 2026-09-01). | A two-platform strategy covers both premium and price-sensitive markets, with the Touch Surgery ecosystem as the data confluence. The rival is Intuitive Surgical's installed da Vinci base. The weakness: whether the two platforms' data interoperate is unstated — and if they do not, the distribution advantage never converts into a training-data advantage. |
| Cornerstone Robotics Hong Kong's in-house surgical robot |
Sentire was developed entirely in-house with a dual-console design, received CE marking in May 2026 and holds approvals in China and Singapore; the company runs three global R&D hubs and six business centres (Medtronic, 2026-09-01). | The moat is regulatory progress and cost structure in Asian markets, plus global distribution acquired through the Medtronic tie-up. The weakness is handing ex-US distribution to a company that simultaneously sells a competing platform of its own, which caps long-run bargaining power. |
| Elucid AI plaque characterisation in cardiovascular imaging |
Announced a $55M Series D on 9/02, taking total funding to about $185M, with investors including an unnamed large public medtech company, IAG Capital Partners and Elevage Medical Technologies; Plaque-IQ is FDA cleared while its FFR-CT version remains under review (MobiHealthNews, 2026-09-02). | Its link to today's through-line: plaque composition analysis is one of the applications spectral photon-counting data most directly enables. The moat is FDA clearance and embedding in existing CT angiography workflows, against HeartFlow and Cleerly. The absence of disclosed outcome data is a conspicuous gap. |
Today's competitive structure is stacked, not head-to-head. The model layer (OpenAI, Google) is fighting over what "reasoning" should look like; the hardware layer (GE, Stryker, Medtronic, Cornerstone) is fighting over who owns the point where data is generated. What is genuinely scarce is neither parameters nor robotic arms but an auditable reasoning record on one side and a real clinical data stream fit for training on the other — and across today's seven items, no single company holds both.
04 — Taiwan Angle
(1) Hallucination and prompt injection stopped being abstract clauses. On May 29, 2026, Taiwan's Ministry of Health and Welfare issued its guidance on generative AI in medical institutions, naming six risks — two of which, "AI hallucination producing false information" and "prompt injection attacks threatening cybersecurity," acquired concrete numbers this week: the BCI paper measured 28.4% of high-confidence outputs as unfaithful, and Astra demonstrated a model finding and exploiting zero-days on its own. The guidance is advisory rather than binding; its nine requirements span pre-deployment assignment of responsibility and data reliability assessment, integration-phase compatibility testing against EHR and imaging systems plus clinical safety validation, and post-deployment retention of final clinical authority by medical professionals, output quality monitoring, bias management and informing patients when AI participates in their care (GlobalBio Monthly).
(2) The three MOHW AI centres map neatly onto what each of today's stories is missing. Taiwan operates a Center for Responsible AI in Healthcare, a Center for External AI Validation in Healthcare (linking cross-hospital data via FHIR and running a federated learning platform) and a Center for Clinical AI Impact Evaluation (health economics analysis and National Health Insurance pricing) (Taiwan's three smart healthcare centres). Line today's cases up against them: an agent like ScreenAgent lacks the external validation the second centre exists to perform; Gemini 3.8 Flash lacks the cost-effectiveness analysis the third would run; and Astra lacks exactly the auditability — regular evaluation, a completed data science cycle — that the first is built around. Three gaps, three centres, and not one of their processes was designed for a model version that changes every three weeks.
(3) Photon-counting CT is the procurement decision Taiwanese hospitals will face soonest. With FDA, Japanese and CE approvals all in hand, Photonova Spectra reaching Taiwan is a matter of timing. But there is an easily overlooked knock-on cost: spectral data is not formatted like conventional CT, so the in-house AI models a hospital has already trained and validated on conventional CT images cannot simply be pointed at the new scanner's output. Buying a photon-counting CT therefore also means sending a batch of already-validated in-house models back for revalidation — a cost that belongs in the procurement assessment, not discovered once the queue at the external validation centre has formed.
05 — Further Reading
-
Intent Drift in LLM-Assisted BCI Communication: An In-Silico Benchmark Under Simulated Decoder Corruption — medRxiv (2026-08-27)
The one item today worth reading in full: it measures "confidence scores cannot serve as a gate" to a precision you could write into a deployment spec, and the method transfers directly to faithfulness evaluation for ambient scribes.
-
Apollo: A multimodal and temporal foundation model for virtual patient representations at healthcare system scale — arXiv (2026-04-20)
Faisal Mahmood's group trained on 25 billion records from 7.2 million patients across 28 modalities and evaluated on 1.4 million patients over 322 prediction tasks. The most complete reference point available for what a record-scale foundation model actually looks like.
-
Learning from routine health system data builds better neuroimaging AI models — Nature Medicine (2026-07-31)
A visual foundation model trained on 5.24 million routine clinical CT and MRI series outperformed models trained on public internet and medical datasets — an answer running against the common assumption that messy in-house data is a liability.
-
OpenAI Astra and Looped Transformers — Sebastian Raschka (2026-09)
Amid a wave of "Astra is hiding its reasoning" takes, this explains what layer looping actually does and argues the opposite. Read it before accepting today's through-line at face value.
-
MedGemma 1.5 model card — Google Health AI Developer Foundations (2026-01-13)
If today's through-line makes you uneasy about frontier model auditability, this is the alternative you can deploy in-house and check benchmark by benchmark; the jump in EHR understanding from 67.6% to 89.6% is worth reading alongside the terms of use.
06 — References
- OpenAI's Astra model is on the way — and very good at breaking into computer systems. TechCrunch, 2026-09-01. techcrunch.com
- OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions. The Decoder, 2026-08. the-decoder.com
- OpenAI Astra and Looped Transformers. Sebastian Raschka, 2026-09. sebastianraschka.com
- Google has released Gemini 3.8 Flash, its fourth Flash model in under four months. Artificial Analysis, 2026-09-02. artificialanalysis.ai
- Google launches two Gemini 3.8 models with cutting-edge reasoning capabilities. SiliconANGLE, 2026-09-02. siliconangle.com
- Gorenshtein A, Omar M, Jia E, Adiniaev Y, Daniel O, Kruskal J, Ahmed M, Brook O, Klang E, Barash Y. Intent Drift in LLM-Assisted BCI Communication: An In-Silico Benchmark Under Simulated Decoder Corruption. medRxiv, 2026-08-27. medrxiv.org
- AI Screens 200,000 Medical Papers for Pennies (ScreenAgent). Yesil Science, 2026-08-30. yesilscience.com
- The Health AI Brief — Week of August 31, 2026. Yesil Science, 2026-08-31. yesilscience.com
- GE HealthCare secures CE mark for its Photonova Spectra technology. MobiHealthNews, 2026-08-31. mobihealthnews.com
- FDA Clears GE HealthCare's MIM Contour ProtégéAI+ 2.0 for Advanced Cancer Treatment Planning. Applied Radiation Oncology, 2026-06-08. appliedradiationoncology.com
- Stryker introduces first of its kind FDA-authorized surgical application for Apple Vision Pro with hip arthroscopy case. Stryker Newsroom, 2026-09-01. stryker.com
- FDA clears Stryker's 'surgical cockpit' that uses Apple's VR headset, Vision Pro. STAT News, 2026-09-01. statnews.com
- Medtronic Announces Strategic Partnership with Cornerstone Robotics to Further Expand Global Access to Robotic-Assisted Surgery. Medtronic Newsroom, 2026-09-01. news.medtronic.com
- Medtronic invests $700M in Cornerstone Robotics. MobiHealthNews, 2026-09-01. mobihealthnews.com
- Elucid raises $55M for AI cardiovascular imaging software. MobiHealthNews, 2026-09-02. mobihealthnews.com
- MedGemma 1.5 model card. Google Health AI Developer Foundations, 2026-01-13. developers.google.com
- Apollo: A multimodal and temporal foundation model for virtual patient representations at healthcare system scale. arXiv:2604.18570, 2026-04-20. arxiv.org
- Learning from routine health system data builds better neuroimaging AI models. Nature Medicine 32(8), 2026-07-31. nature.com
- 衛福部首發醫療 GenAI 指引:六大風險、九項要點建立治理框架. 環球生技月刊, 2026-05-29. news.gbimonthly.com
- 臺灣智慧醫療三大中心. 衛生福利部. aicenter.mohw.gov.tw
- AI model release log (Gemini 3.8 Flash, Claude Fable 5.1, GLM-5.3-Flash, Qwen3.8 Flash, Granite 4.2). LLM-Stats, 2026-09. llm-stats.com