Restricted analysis, made public daily.
Declassified under standing order Edition No. 073 Monday, September 14, 2026

The Question Never Changed.
Claude’s Answer Did.

Peer-reviewed research tested three AI models on the same benchmark questions, attaching different user biographies. Accuracy fell for all three in some conditions. One model, Claude 3 Opus, also stood out for refusing to answer, and for condescending language toward lower-status personas.

The question never changed. The biography attached to it did. Peer-reviewed research presented at AAAI 2026 tested GPT-4, Claude 3 Opus, and Llama 3-8B on 817 TruthfulQA and 1,000 SciQ questions, prompting each with a persona biography — education level, English proficiency, and country of origin, in varying combinations — and comparing the results against a control condition with no biography at all, across four runs each. All three models’ accuracy fell in at least some of those conditions. Claude 3 Opus stood out in its refusal behavior.

That model is a Claude model, specifically the Claude 3 Opus checkpoint. Its refusal rate reached 10.9% for the foreign, low-education group, averaged across datasets — about three times its own 3.61% rate with no biography at all. In a separate, manual analysis of its refusals specifically, the authors found condescending or patronizing language — unprompted, simplified “broken” English — in 43.74% of Claude’s refusals toward less-educated personas, versus under 1% toward highly-educated ones, and under 1% for the other two models. On TruthfulQA, Claude scored 66.22% for the Iran, low-education persona, against 78.17% with no biography at all.

The benchmark didn’t move. Three separate measures of Claude did.

What moved wasn’t the facts

The multiple-choice question never changed between runs. Only the biography attached to it did. What the study measured is model behavior on a fixed benchmark, run four times against 817 and 1,000 questions — not any real person’s field of work, credentials, or professional standing. That distinction matters for what the finding can honestly be asked to carry.

The Agentics read: a question for authority infrastructure

This study didn’t test professional recognition. It tested benchmark accuracy and refusal behavior against biographies varying education level, English proficiency, and country of origin, compared to a no-biography control. Accuracy fell for all three models in some conditions; Claude 3 Opus additionally stood out for refusals and condescending language. Whether the same pattern reaches how an AI system represents a real professional’s actual field of work is a question this study raises, not one it answers. Contextual Ambiguity is Agentics’ own interpretive lens on that open question — a framework this publication brings to the record, not a diagnosis the researchers themselves made.

Editorial disclosure: Claude is used in this publication’s production process. The study tested the February 2024 Claude 3 Opus checkpoint; its findings do not automatically describe the Claude system used here.

Sources

Poole-Dayan, E., Roy, D., & Kabbara, J., “LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users,” arXiv:2406.17737, tested GPT-4 (gpt-4-0125-preview), Claude 3 Opus (claude-3-opus-20240229), and Llama 3-8B; now published in the peer-reviewed proceedings of AAAI 2026 (DOI 10.1609/aaai.v40i46.41259, Proceedings of the AAAI Conference on Artificial Intelligence 40(46):39116–39124): abstract, full text, AAAI record.

MIT News covered the underlying research on model performance disparities across user profiles: MIT News.

How does AI represent your standing
when the context changes?

Bring that question to SIA.

Every edition, in order, from No. 001 · A new edition releases daily, 05:30 CT.

Open the Record
‹ Edition No. 072 Edition No. 074 ›