Restricted analysis, made public daily.
Declassified under standing order Edition No. 053 Tuesday, August 25, 2026

The Hallucination Rate Doubled.

The Test Disappeared.

OpenAI built a benchmark that measured one thing: whether its models get the facts right about real, named people. Beginning in December 2024, it appeared repeatedly in the company’s system and model cards. In April 2025 it posted a reading that, by the card’s own words, needed more research to understand. Its last appearance in the OpenAI system and model cards this publication checked was August 5, 2025. Eighteen cards have followed. It appears in none of them.

In December 2024, a line item called PersonQA appeared in the system card for OpenAI’s o1 — absent from the GPT-4o and o1-preview cards that preceded it. The card defined it plainly: a dataset of “questions and publicly available facts about people that measures the model’s accuracy on attempted answers.” Not facts in general. Facts about people. Of all the instruments OpenAI could publish, this was the one built to answer the question every professional quietly asks: how often is it wrong about a real person? The benchmark ran again in the o3-mini and Deep Research cards that winter. In February, the GPT-4.5 card set what remains the benchmark’s accuracy high-water mark for a model answering without browsing: 0.78, with a hallucination rate of 0.19. The Deep Research card, published the same month, returned better numbers on both measures, by its own count (0.86, and a 0.13 the card itself calls an overstatement of its true error rate) — but Deep Research answers by searching, and the card credits “the heavy reliance on online search” with reducing exactly these errors.

Then came the April 2025 card for o3 and o4-mini, and the instrument did its job. OpenAI’s new leading reasoning model, o3, posted a PersonQA hallucination rate of 0.33 — roughly double the o1 comparison value of 0.16. (The smaller o4-mini logged 0.48, three times the o1 comparison value, which the card calls expected: smaller models carry less world knowledge and hallucinate more.) The card is careful and candid about what it saw: o3 “tends to make more claims overall, leading to more accurate claims as well as more inaccurate/hallucinated claims.” Its accuracy did rise, 0.59 against o1’s 0.47. The model had not gotten dumber about people. It had gotten more willing — right more often, and wrong far more often, at the same time. And the card names where the wrongness concentrated: “While this effect appears minor in the SimpleQA results, it is more pronounced in the PersonQA evaluation. More research is needed to understand the cause of these results.”

The instrument itself outlived the spring. The ChatGPT Agent card (July 2025) ran PersonQA with browsing enabled and reported o3’s hallucination rate at 0.024, noting its results were “better than those in the o3 system card, because the metrics in the o3 card reflected performance without browsing.” The gpt-oss model card ran it once more for the open-weight models on August 5, 2025, making the same point in general terms: browsing reduces hallucination because models “are able to look up information they do not have internal knowledge of.” The gpt-oss run is the benchmark’s last appearance in any card reviewed for this edition.

Days later, the GPT-5 system card arrived with its hallucination section intact — SimpleQA and other instruments still present — and PersonQA absent. To confirm the picture held, this publication saved and checked the full dated listing of OpenAI’s Deployment Safety Hub on August 24, 2026 — the census below.

In the eighteen cards published after August 5, 2025, the GPT-5 card among them, the benchmark appears exactly zero times. The company’s own April card had posed the question of the o3 regression; nothing in the set ever answered it.

Something did replace it, though not everywhere (several cards cover products with no factuality section at all), and not with the same question. The GPT-5 card, the first without PersonQA, grades factuality on prompts representative of real ChatGPT production conversations, scored by a model-based grader with web access. The current flagship-family cards evaluate factuality against user-flagged conversations; the GPT-5.5 and GPT-5.6 (Sol) cards put it verbatim — “de-identified ChatGPT conversations that users of our prior models have flagged as containing factual errors,” cases those cards describe as especially hallucination-prone, “not a representative slice of all production traffic.” Real measurement, arguably harder grading. But it measures what users complained about. Nothing in either method isolates the question PersonQA existed to isolate: whether the machine gets real, named people right.

0.33 against 0.16, flagged by its own card as unresolved — and never revisited. Since the benchmark’s last appearance in August 2025, eighteen OpenAI cards have followed; it appears in none.
The census — 28 OpenAI documents, dated August 24, 2026

Reviewed: the 22 card entries in the Deployment Safety Hub’s own saved listing, plus 6 earlier cdn-hosted cards — August 2024 through August 2026.

PersonQA present — 7: o1 (Dec 5, 2024) · o3-mini (Jan 31, 2025) · Deep Research (Feb 25, 2025) · GPT-4.5 (Feb 27, 2025) · o3 & o4-mini (Apr 16, 2025) · ChatGPT Agent (Jul 17, 2025; browsing-enabled) · gpt-oss (Aug 5, 2025 — final appearance).

PersonQA absent — 21. Never carried it (3): GPT-4o (Aug 2024) and o1-preview (Sept 2024), which predate the benchmark, and the GPT-4o image-generation addendum (Mar 2025), which has no factuality section. After its last appearance (18): every hub card published after August 5, 2025, from GPT-5 (Aug 2025) through GPT-5.6 (Sol; Jul 9, 2026) and its August 6, 2026 update.

What the record says, and what it doesn’t

That reading is real. What it licenses is narrower than it feels, so here are the boundaries, stated plainly: benchmarks get replaced all the time (they saturate, they leak into training data, better instruments supersede them). OpenAI never promised PersonQA was permanent, never explained its absence, and owes no one an explanation for routine engineering churn. Nothing in the public record supports any claim about why it stopped appearing, and this edition makes none. Nor did the benchmark vanish the moment it embarrassed anyone: it ran twice more, in other product lines, before it stopped. The April card itself even notes that comparison values from live models “may vary slightly from values published at launch” — o1’s own PersonQA number read 0.20 in February and 0.16 in April under the same label. The instrument was imperfect and the baseline unstable. The disappearance has no stated cause — and is unremarkable in kind.

The claim that survives every fair reading is harder to escape: the o3 regression has stood unexplained for sixteen months, and for the last twelve months (eighteen cards) the instrument that measured it has not appeared at all. The systems kept shipping, and the answers about people kept flowing. The gauge went dark.

Anatomy of an instrument that stopped appearing

The instrument was alone in its class. Among the hallucination evaluations OpenAI’s cards name — general-facts sets, production-conversation graders, flagged-complaint samples — PersonQA was the only one whose subject was real, named people.

The o3 reading was never superseded — no later card names the 0.33 figure, re-runs the comparison, or replaces it with a comparable non-browsing measure.

The successor methods measure complaints and production traffic, not persons. They are legitimate instruments. They answer a different question. Filed as an Agentics Authority Misclassification specimen, one of the five clinical characteristics of Digital Derangement Syndrome: person-specific accuracy stopped being isolated in a published benchmark where anyone can read the result.

The question that outlived the test

If your name, your license, your firm, or your record is an asset — and for anyone whose clients arrive by reputation, it is the asset — OpenAI’s current published system cards provide no person-specific benchmark result. Nothing in them tells a professional how reliably these systems represent real, named people. The only OpenAI benchmark in this review set designed specifically around questions and public facts about people has not appeared in a published card for more than a year — and the question its April reading raised is still open. Even a person-specific benchmark, measured in the aggregate, would not tell you whether the system gets your particular record right.

The reading has to come from your own record instead: what the machines actually return about you, checked against what is true, backed by identity infrastructure the machines can verify rather than guess at. OpenAI attributed the far lower browsing-enabled PersonQA rate in the ChatGPT Agent evaluation to the browsing configuration itself — and the cards do not identify which sources produced that difference or establish that browsing will improve every person-specific answer. That is the ground Answer Engine Authority occupies: making a professional’s public record clearer, more coherent, and easier to verify when an answer system consults external sources. It cannot guarantee what the system will retrieve, believe, or say. It can improve the evidence available to be found.

Run the one-line audit yourself: ask which benchmark in OpenAI’s current published system cards specifically measures accuracy about real, named people. Since August 5, 2025, across the twenty-eight documents in the census above, the answer is: none.

Sources

All quoted language and figures are from OpenAI’s own published system cards: the o1 System Card (December 5, 2024: PersonQA’s earliest appearance among the cards checked for this edition; the GPT-4o and o1-preview cards that precede it carry no mention), the o3-mini System Card (January 31, 2025), the Deep Research System Card (February 25, 2025), the GPT-4.5 System Card (February 27, 2025: PersonQA accuracy 0.78, hallucination rate 0.19; its o1 comparison column reads 0.20), the o3 and o4-mini System Card (April 16, 2025; Table 4: PersonQA hallucination rate o3 0.33, o4-mini 0.48, o1 0.16; accuracy o3 0.59, o1 0.47; “More research is needed to understand the cause of these results”), the ChatGPT Agent System Card (July 17, 2025: PersonQA run with browsing enabled; o3’s rate 0.024, results “better than those in the o3 system card, because the metrics in the o3 card reflected performance without browsing”), and the gpt-oss Model Card (August 5, 2025: PersonQA’s last appearance; “browsing or gathering external information tends to reduce instances of hallucination as models are able to look up information they do not have internal knowledge of”).

The absence claim is bounded and was checked on August 24, 2026, against a dated roster of OpenAI’s Deployment Safety Hub: the twenty-two published card entries its own index listed that day, Deep Research (February 2025) through the GPT-5.6 August update (August 6, 2026), saved at capture. PersonQA appears in seven checked documents, all named above, and in none of the eighteen hub cards published after August 5, 2025, including the GPT-5 System Card (cover-dated August 13, 2025, listed by the hub as published August 7; verified against the document’s extracted text, which does still name SimpleQA; that card grades factuality on prompts representative of production conversations) and the GPT-5.6 (Sol) card (July 9, 2026; a distinct document from the August 6 update that closes the roster). The GPT-5.5, GPT-5.6 Preview, and GPT-5.6 (Sol) cards describe, verbatim, evaluating “de-identified ChatGPT conversations that users of our prior models have flagged as containing factual errors”; the family’s remaining cards state the same method with minor wording variants. Several hub cards cover non-chat products (image, video, voice, and research previews); the GPT-5 and 5.5/5.6-family cards carry the real weight of the check. This edition asserts nothing about documents outside that checked set, and nothing about why the benchmark stopped appearing.

The Test Stopped Appearing.

You’re Still Being Graded.

Every day, answer systems assert claims about real people to the people deciding whom to trust. The one person-specific instrument in the OpenAI cards reviewed for this edition posted a doubled hallucination rate, a result the company said required more research to understand — then stopped appearing in the eighteen cards that followed.

You cannot subpoena a benchmark that no longer publishes. Your own record is the instrument now: what the machines return about you, compared with the documented record, then traced to the public sources that verify it.

How that reading runs for a working professional is a question for SIA — the Intelligence Officer, briefed on
every edition of this record the morning it releases.

Every edition, in order, from No. 001 · A new edition releases daily, 05:30 CT.

Open the Record
‹ Edition No. 052 Edition No. 054 ›