The Prompt Changed
Who Made the List.
Researchers tested forty-three AI models across variations in persona and request. The composition of the recommended scholar lists changed with the prompt.
A May 2026 preprint audited forty-three large language models using a shared scholar-recommendation template. Researchers varied the assigned persona’s role, location, and language, alongside the requested discipline, subfield, seniority, and number of recommendations. The audit repeated this template across six disciplines: biology, computer science, mathematics, physics, psychology, and sociology. The researchers assessed technical quality and representativeness against bibliographic reference data from Semantic Scholar and OpenAlex, using name-inferred demographic proxies. The paper’s own summary of what it found: “the model primarily determines whether responses are well-formed, whereas the prompt primarily determines who gets recommended.”
Model choice mattered most for basic response quality. Request context — the field, the subfield, the seniority level, how many names were asked for — generally mattered more than persona for recommendation outcomes, with location the important exception. Its clearest demonstration of that exception sits in a single contrast: “South Africa prompts yield less factual lists, while Japan prompts yield highly factual but homogeneous lists skewed toward highly productive scholars.” These are aggregate patterns across prompt configurations, each summarizing many runs rather than one list held against another. Homogeneity alone is a signal worth checking further. What matters is whether the selection fits the request.
Recognition, filtered before it starts
The audit varied the request, not the scholars’ credentials. Its findings concern how prompt composition affects recommendations. Whether anyone became more deserving between runs is a separate question entirely. A recommendation speaks to one context. Expertise is bigger than that.
Different recommendations can be appropriate to different needs. A search for junior researchers and a search for a senior collaborator reasonably need different lists; so do a PhD student seeking an advisor and a recruiter seeking hires, and geographic relevance can be legitimate too. The concern is whether the answer remains accurate and whether its selection criteria are clear. Agentics interprets this context-sensitive recognition through Digital Derangement Syndrome™’s Contextual Ambiguity characteristic — a framework this publication brings to the record, separate from the study’s own conclusions.
The study measures recommendations, not hiring. Forty-three models were asked for scholar recommendations and checked against a database of who publishes in a field. The record holds only what the system said when asked.
Posture, precisely. The May 27, 2026 arXiv record for “Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation” (Sánchez-Guzmán, Eberhard, Helic & Espín-Noboa, arXiv:2605.28187) lists it as under review. This edition evaluates that preprint; its predecessor’s KDD ’26 proceedings status does not confer peer-reviewed status on its successor.
The record with no ballot in it
This raises a question central to Answer Engine Authority™: under which requests does expertise become visible, and does the recommendation fit the evidence and the request? One appearance is a data point, not a ballot; recognition has to be examined across the situations that matter. Becoming microfamous means being known where it counts — including when the relevant request reaches an AI system.
Sánchez-Guzmán, A., Eberhard, L., Helic, D., & Espín-Noboa, L., “Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation,” arXiv:2605.28187 (submitted May 27, 2026; arXiv record lists “under-review”): abstract, full text.
Espín-Noboa, L. & Méndez, G.G., “Whose Name Comes Up? II,” arXiv:2602.08873 — in the proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’26); cited for research-lineage context, its findings not attributed to the newer, unreviewed paper above: abstract.
Semantic Scholar (semanticscholar.org) and OpenAlex (openalex.org) — the bibliographic reference databases, together with name-inferred demographic proxies, used to assess factuality and representativeness of the AI-recommended scholar lists.