Restricted analysis, made public daily.
Declassified under standing order Edition No. 024 Monday, July 27, 2026

Nature Medicine tested it.
Most engines never mentioned it.

A Mount Sinai team put ChatGPT Health’s triage advice through a structured test — 960 responses, with the failures concentrated exactly where the stakes run highest — and published what they found online in Nature Medicine on February 23rd. Five months later, OpenAI opened the product to every American adult with an account. So on July 26th we asked six answer engines the simplest question there is about it, and four of them answered as if the study had never been written.

On July 23rd, OpenAI opened Health in ChatGPT to every logged-in adult in the United States, on every tier including the free one. The same change extended connected health context into ordinary conversation — the company found that more than 70% of health-related conversations were happening outside the space built for them.

Five months earlier, on February 23rd, Nature Medicine had published a structured examination of the product by a team at the Icahn School of Medicine at Mount Sinai. Sixty clinician-authored vignettes across twenty-one clinical domains, each run under sixteen systematic variations, producing 960 responses. The authors reported a system strongest in the middle and weakest at the extremes, with the failures concentrated at the two ends where the stakes are highest: error rates of 35% on nonurgent presentations and 48% on emergency conditions. Among cases the authors classed as gold-standard emergencies, the system undertriaged 52% — routing diabetic ketoacidosis and impending respiratory failure toward evaluation within 24 to 48 hours rather than to an emergency department — while, the authors note, correctly triaging classical emergencies such as stroke and anaphylaxis. The authors also reported that when family or friends in the vignette minimized the symptoms, triage recommendations shifted significantly in edge cases — an anchoring effect they quantified in the paper.

Whether that is disqualifying, or the expected performance of a consumer tool operating under a disclaimer, is a question for clinicians and regulators. This edition takes no position on it. The vendor has published its own account of this product. A research team has published a test of it. What follows is only about which one travels.

On July 26th we put a single identical question — What is ChatGPT Health? — to six answer engines in clean sessions — no prior conversation in the thread, consumer accounts, United States: ChatGPT, Gemini, Perplexity, Google’s AI Overview, Microsoft Copilot, and Claude. One run each on one day, except a single engine run twice — its first pass failed to resolve the product at all — and the study surfaced in neither pass. An observation, then, not a study. Two of the six are Google surfaces, and Copilot runs partly on OpenAI models — so the stricter count is four independent stacks, one of which was asked to describe its own maker’s product. All six, by the end, resolved the product correctly. They returned the connected data sources, the tiers, the availability, the privacy architecture, the disclaimer. Four of the six turned up no independent safety finding whatsoever — the engine describing its own maker’s product among them. A fifth gestured at “peer-reviewed studies” and “journals like Nature Medicine,” naming no study, no authors and no figure, against a bare link to the journal’s front door. One named the paper and its numbers — and disclosed, unprompted, that it was built by a competitor.

The study was five months old, in one of the most cited journals in medicine, with a DOI and a PubMed identifier. Four of six answer engines described the product without it.

What the engines carried instead

Process, performance, and the silence

The vendor publishes process. OpenAI’s account of Health describes a global network of more than 260 physicians across 60 countries, 49 languages and 26 medical specialties, who have reviewed more than 700,000 example model responses. Every figure is a measure of review performed — an input.

The researchers published performance. The Mount Sinai study measured what the system did when tested — an output. The two are not in conflict about facts. They answer different questions, and only one of them is a result.

Four of six carried neither. A reader asking a direct question about this product received its description — sources, tiers, availability, privacy, disclaimer — and no independent result at all. Not a distortion, not a refusal. The finding simply was not there.

Publication is not retrieval

There is a comfortable assumption underneath most professional strategy: that rigor eventually surfaces. Do the work, submit it, survive review, and the record corrects itself. Peer review is the most expensive verification we perform, and the assumption is that engines built on the written record inherit the result.

They do not inherit it. Retrieval is its own system, with its own inputs: entity structure, corroboration, freshness, the shape of the question asked. A journal article is a document; whether it is retrievable as the answer to a given question is a separate property, one that peer review does not confer and impact factor does not guarantee. A vendor’s own page about its own product, meanwhile, is exactly what a question phrased as what is this product is built to return.

The obvious reply is that we asked a definitional question and are complaining about getting a definition. But the person asking what is this is precisely the person who has not yet decided, has no specialist literature to fall back on, and most needs to know the thing was independently tested. That is the question a lay decision-maker actually asks, and answering that person without the finding is the defect. And the study is not missing from the indexes — later the same day, handed a question that already carried its findings, one of the four returned the paper, Mount Sinai’s own announcement, and the PubMed record. The record is there. The definitional question never touches it.

The finding does not stop at medicine. If a February paper in Nature Medicine, from a named team at Mount Sinai, cannot reliably reach a reader asking a direct question about the very product it examined, then the white paper, the case study and the decade of results sitting behind your own claims are not losing to a better argument — they are simply not retrievable at the moment the question is asked.

The record is not self-correcting. It is retrieved, or it is not, and something decides which. That thing ran today, in every room where somebody asked, and nobody submitted anything to it for review.

Sources

Ramaswamy A, Tyagi A, Hugo H, et al. “ChatGPT Health performance in a structured test of triage recommendations.” Nature Medicine 32(5):1671–1675, published online February 23, 2026 (print issue May 2026) — DOI 10.1038/s41591-026-04297-7, PMID 41731097; institutional summary at Mount Sinai’s newsroom. The 35% and 48% figures are error rates for nonurgent presentations and emergency conditions respectively; the 52% figure is the undertriage rate among cases the authors classed as gold-standard emergencies. These are three distinct measures and are not interchangeable.

Health in ChatGPT was announced January 7, 2026, and extended to all logged-in U.S. users aged 18 and over, across Free, Go, Plus and Pro, on July 23, 2026, at which point connected health context also became available in general conversation — OpenAI, reported by TechCrunch. The physician-network figures are OpenAI’s own, published at Improving health intelligence in ChatGPT.

The six-engine result is our own, gathered July 26, 2026, using one identical prompt in clean sessions, and is reported here as a dated observation rather than a controlled study. This edition takes no position on the clinical safety of the product examined. This publication’s production tooling runs on one of the engines tested.

The paper answered for it.
Nothing answers for you.

ChatGPT Health had something almost no business will ever have: a named research team at Mount Sinai examined it and filed the result in the permanent record — and even that barely surfaced. In most markets nobody is independently establishing what you do, publishing it where machines look, or structuring it so a direct question returns you. That work is installed, never accumulated — entity architecture, corroboration, structure — and the installation is Answer Engine Authority. Something already answers for Agentics: her name is SIA, and she is what an installed answer sounds like. The engines answered a health question six times on Sunday. They are answering yours, in your market, today — from whatever happens to be retrievable.

Every edition, in order, from No. 001 · A new edition releases daily, 05:30 CT.

Open the Record
‹ Edition No. 023 Edition No. 025 ›