Restricted analysis, made public daily.
Declassified under standing order Edition No. 065 Sunday, September 6, 2026

The Citation Records the Credit.
Not the Reason.

A study published in this year’s ACM SIGIR proceedings ran 252,000 controlled trials across six AI models to isolate one question: when two sources compete to be cited, what decides the winner? Two drivers above all others — the topic the source matches, and the slot it happened to occupy. Only one of them is written on the page.

The experiment was built to remove every excuse. Two candidate sources, injected directly into an AI model’s context, identical in every respect but one. Brands anonymized, so reputation could not vote. Presentation order counterbalanced, so the test could separate what a source says from where it sits. Then one measurement. Which of the two does the model’s first citation credit? The authors ran that trial 252,000 times, across eighteen content factors and six models: Gemini 2.5 Flash, Claude 3.5 Sonnet, GPT-5.2, GPT-5 Mini, GPT-5 Nano, and Kimi K2 Thinking. Then they let mixed-effects models say what actually decided. The verdict: “topical relevance and list position are the biggest drivers of being cited first.”

Read that pairing again. One of the two biggest drivers is about the source. The other is about the seat it was given. The paper sorts its eighteen factors into a hierarchy, and at the top it names four recurring conditions its authors call gatekeepers: topical match, explicit price information, a recent timestamp, and list position. One caution belongs on the record beside them: the paper’s summary sentence claims uniformly large effects for all four across all six models, but its own Table 2 is less tidy: the price row runs from single digits on some models to past ten thousand on others, and the timestamp row from the teens to past ten thousand. The row this edition rests on is position: very strong, by the paper’s own scale, on all six models. Three of those four conditions are things a source can write: match the topic, state the price, carry a current date. The fourth is not written anywhere on the page. In this experiment, the researchers assigned the slot. How production systems assign it is a separate question.

The magnitudes deserve careful language, and the record will use the paper’s. On the position-one-versus-position-two comparison, four of the six models registered odds ratios above 10,000; the two most position-tolerant came in at 2,002 and 1,795. These are fitted model estimates, not observed win counts. Where preference ran nearly deterministic, the authors say so themselves, reporting extreme values as “>10k where applicable” and treating them as “indicating a decisive win for variant A rather than a finely resolved numeric ratio.” The Gemini position estimate carries a degenerate-Hessian convergence warning, the paper’s own flag. Read with that discipline, the honest summary is still striking: the authors describe two models, Gemini and Claude, as categorical, with 67–78% of their significant factors producing effects above 10,000, “suggesting binary decision boundaries.” Within this testbed, position strongly favored one source receiving the first citation. That does not establish that the other source went unread or uncited (in roughly one successful answer in ten, the models cited more than one URL). But the first credit, which a reader may interpret as a signal of authority, followed the seat.

Both sources were already in the room. Their order strongly influenced which received the first credit.

What the testbed can say. What it cannot.

This record files findings with their boundaries attached, and this study states its own. It is a two-document testbed, not a live answer engine: “Production RAG often retrieves five to ten or more pages, so real citation pools are larger than our testbed.” And the controls that make the isolation clean also remove forces that operate in the wild: “Production systems may still favor trusted domains or strong brands when they pick sources.” So the claim on file is not that your last AI answer ran this experiment. The claim is narrower, and harder. With content and position separated under controls, both topical relevance and position strongly influenced the first credit — and only one of them is on the page. Nor does the phenomenon stand alone. Position-weighting in long contexts was established in Lost in the Middle two years before this paper measured it at the citation event itself. One more line belongs in the open. All three authors are employees of Sprinklr, a customer-experience software company, an affiliation the paper itself discloses along with an internal pilot of the derived tooling. The ACM proceedings venue is on the record; the affiliation rides beside it anyway, because that is what a record is for.

Anatomy of a citation influenced by position

What was tested. Two sources, one factor apart, injected into six models’ contexts 252,000 times; brands anonymized, order counterbalanced. Outcome: which source the first citation marker credits.

What drove the credit. Topical relevance — and list position. The paper names four recurring gatekeeper factors (topic match, a stated price, a recent timestamp, the slot), with per-model effects its own Table 2 shows varying widely for price and timestamp. Position one versus position two registered fitted odds ratios above 10,000 on four of six models. On two of those, Gemini and Claude, the paper’s own reading of the wider pattern is “binary decision boundaries.”

What that means clinically. A citation can be read downstream as a verdict on authority — who was worth crediting. This instrument shows the first credit tracking, in controlled conditions, substantially with position. In this record’s clinical vocabulary — the mapping is ours, not the paper’s — that is Authority Misclassification: standing conferred on a signal that is not authority. It also raises, as a downstream implication to test rather than an experimentally established diagnosis, Decision Exclusion: if the visible credit follows the seat at these magnitudes, the source one slot down can lose the reader’s attention with no trace of the contest in the answer.

This record has filed the thesis before the mechanism. Edition No. 007 put it in one sentence: you will not be turned down; your name will simply not come up. This study does not prove that sentence — it measures something adjacent and quieter: when two names are both in the room, which one gets credited first, and how steeply that follows the seat. It is also the second time this record has filed research from this year’s SIGIR proceedings: Edition No. 049 documented a different paper showing that cosmetic differences in how the same question is typed — punctuation included — change which sources are returned. Between the two, the shape of the pipeline is on the record: the query shapes which sources come back, and the seat leans hard on which one is credited first.

Why this reaches past the lab

The professionals this record is written for do not think of themselves as documents in a retrieval pool. Their published work can become a candidate whenever an answer engine gathers sources relevant to their expertise. What this paper hands them is not a prescription — its scenarios were product comparisons, its controls deliberately unlike production — but a set of questions worth testing against a real practice. Does the record match the questions actually asked? Are the concrete facts a machine can check — a price, a date — stated where it looks? And the uncomfortable one: if the first credit follows the seat at any fraction of what the testbed measured, what assigns the seat? That makes the ordering process worth investigating alongside the content. This experiment does not establish how a production system assigns that order — or which changes would improve it. What it establishes is that position is worth interrogating at all — at magnitudes steep enough to make the interrogation urgent.

And that is the honest close. The citation beneath an AI answer can look like a judgment — as though the machine surveyed the field and credited the worthiest source. This edition files the controlled result that says the reading is, at minimum, incomplete. In this comparison, with everything else stripped away, both topical relevance and position strongly influenced who got credited first — and position is not itself evidence of expertise. The next time an answer credits one practitioner over another in your field, the record holds a documented second explanation to weigh. The citation records who received credit. It does not explain why.

Sources

Rahul Vishwakarma, Shushant Kumar & Ratnesh Jamidar, “What Gets Cited: Competitive GEO in AI Answer Engines,” Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’26, Melbourne), Association for Computing Machinery, published July 19, 2026 — DOI 10.1145/3805712.3808445; preprint at arXiv:2605.25517 (submitted May 25, 2026). All quoted passages extracted directly from the paper; venue and publication metadata verified against the ACM DOI record via Crossref, September 5, 2026. Odds-ratio figures are the paper’s own (Table 2, mixed-effects models); entries the paper reports as “>10k” are cited here exactly that way. Convergence note: the paper flags extreme estimates for quasi-separation and marks convergence warnings in its Table 2 legend (degenerate-Hessian and singular-fit flags); its own guidance is to treat “>10k” entries as “indicating a decisive win for variant A rather than a finely resolved numeric ratio.”

Disclosure, filed with the finding: all three authors are Sprinklr employees; the paper discloses an internal pilot and beta deployment of the derived practitioner tooling at Sprinklr. Publication in the ACM SIGIR ’26 proceedings is the venue of record; this record’s rules require the affiliation to appear beside the result regardless.

Nelson F. Liu et al., “Lost in the Middle: How Language Models Use Long Contexts,” Transactions of the Association for Computational Linguistics (2024) — arXiv:2307.03172. Independent prior on position-weighting in long contexts.

Method note: the study measures which of two injected candidate sources an answer engine cites first in a controlled testbed; it does not observe live production ranking, and its authors say so in the limitation lines quoted above.

The next cited answer will name a source.
The citation alone will not explain why that source received credit.

The paper measured the seat’s power under laboratory controls. Whether it holds that power over your record —
in the pools where real answers get assembled — is
a question for SIA — the Intelligence Officer,
briefed on every edition of this record
the morning it releases.

Every edition, in order, from No. 001 · A new edition releases daily, 05:30 CT.

Open the Record
‹ Edition No. 064 Edition No. 066 ›