The Question Never Changed.
Claude’s Answer Did.
Peer-reviewed research tested three AI models on the same benchmark questions, attaching different user biographies. Accuracy fell for all three in some conditions. One model, Claude 3 Opus, also stood out for refusing to answer, and for condescending language toward lower-status personas.
The question never changed. The biography attached to it did. Peer-reviewed research presented at AAAI 2026 tested GPT-4, Claude 3 Opus, and Llama 3-8B on 817 TruthfulQA and 1,000 SciQ questions, prompting each with a persona biography — education level, English proficiency, and country of origin, in varying combinations — and comparing the results against a control condition with no biography at all, across four runs each. All three models’ accuracy fell in at least some of those conditions. Claude 3 Opus stood out in its refusal behavior.
That model is a Claude model, specifically the Claude 3 Opus checkpoint. Its refusal rate reached 10.9% for the foreign, low-education group, averaged across datasets — about three times its own 3.61% rate with no biography at all. In a separate, manual analysis of its refusals specifically, the authors found condescending or patronizing language — unprompted, simplified “broken” English — in 43.74% of Claude’s refusals toward less-educated personas, versus under 1% toward highly-educated ones, and under 1% for the other two models. On TruthfulQA, Claude scored 66.22% for the Iran, low-education persona, against 78.17% with no biography at all.
What moved wasn’t the facts
The multiple-choice question never changed between runs. Only the biography attached to it did. What the study measured is model behavior on a fixed benchmark, run four times against 817 and 1,000 questions — not any real person’s field of work, credentials, or professional standing. That distinction matters for what the finding can honestly be asked to carry.
This study didn’t test professional recognition. It tested benchmark accuracy and refusal behavior against biographies varying education level, English proficiency, and country of origin, compared to a no-biography control. Accuracy fell for all three models in some conditions; Claude 3 Opus additionally stood out for refusals and condescending language. Whether the same pattern reaches how an AI system represents a real professional’s actual field of work is a question this study raises, not one it answers. Contextual Ambiguity is Agentics’ own interpretive lens on that open question — a framework this publication brings to the record, not a diagnosis the researchers themselves made.
Editorial disclosure: Claude is used in this publication’s production process. The study tested the February 2024 Claude 3 Opus checkpoint; its findings do not automatically describe the Claude system used here.
Poole-Dayan, E., Roy, D., & Kabbara, J., “LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users,” arXiv:2406.17737, tested GPT-4 (gpt-4-0125-preview), Claude 3 Opus (claude-3-opus-20240229), and Llama 3-8B; now published in the peer-reviewed proceedings of AAAI 2026 (DOI 10.1609/aaai.v40i46.41259, Proceedings of the AAAI Conference on Artificial Intelligence 40(46):39116–39124): abstract, full text, AAAI record.
MIT News covered the underlying research on model performance disparities across user profiles: MIT News.