Restricted analysis, made public daily.
Declassified under standing order Edition No. 059 Monday, August 31, 2026

It Cited a Source.
Nobody Checked What Wrote It.

A Northwestern University audit of ChatGPT, Copilot, Gemini, and Perplexity found that a classifier flagged roughly 16% of the sources these engines cited as likely or highly likely AI-generated — on Copilot, nearly 3 in 10. Of 200 cited pages classified as likely or highly likely AI-generated, none disclosed AI use.

Northwestern University researchers Mowafak Allaham and Nicholas Diakopoulos set out to answer a narrow, mechanical question: when a generative search engine cites a source, does the citation itself carry any signal of what actually produced the content behind it? Their audit ran 712 real, human-submitted queries on politics, health, and the environment through four generative search engines — ChatGPT, Copilot, Gemini, and Perplexity — retrieving and checking 72.9% of the sources those engines cited against an AI-content classifier the researchers chose specifically because it errs conservative, so their own headline number functions as a floor, not a ceiling. Even at that floor, Pangram classified roughly 16% of the 19,154 unique cited URLs the researchers successfully scraped as likely or highly likely AI-generated.

The rate isn’t uniform across engines. Copilot cited AI-generated sources in 27.8% of its citations (N=1,614) — what the paper itself calls “nearly 3 out of every 10 cited sources being AI-generated.” Gemini followed at 14.7% (N=782), Perplexity at 9.4% (N=472), and ChatGPT at 7.3% (N=765). The measured rates differed substantially by provider, although this single audit cannot establish that the same proportions persist over time.

The researchers didn’t stop at the aggregate number. Ranking domains by the number of cited URLs Pangram classified as AI-generated, www.factually.co ranked eighth among the 35 domains with the most such URLs, associated with 28 unique cited URLs Pangram classified as highly likely AI-generated — from a site whose own homepage describes it as an “AI-powered research tool” and adds, unprompted: “Is it perfect? No. Accuracy is hard. Consistency is hard. We’re working on it.” The site isn’t hiding what it is. Across the sample, disclosure was checked directly: the researchers hand-checked 200 URLs the classifier had flagged as likely or highly likely AI-generated for any statement disclosing AI use. They found zero.

A citation is supposed to make a claim inspectable: where did it come from, who published it, and what evidence stands behind it? In this audit, roughly one successfully analyzed source in six contained text Pangram classified as likely or highly likely AI-generated. In the researchers’ 200-page disclosure check, none said so. The link was visible. A consequential part of its provenance was not.

What the Footnote Doesn’t Tell You

Anatomy of undisclosed provenance

The researchers did not take a single detector entirely on faith. On their curated 200-human/200-AI test, Pangram correctly classified all 200 human texts; the paper separately notes that Pangram’s provider reports a false-positive rate of roughly one in 10,000. On a second, 105-article set previously identified by journalists as AI-generated, Pangram and GPTZero agreed on 82% of labels, while Pangram identified 72 of the 105 articles as AI-generated after the researchers reclassified its mixed results. That miss rate is why the authors treat their prevalence estimate as a lower bound — while also acknowledging that a different detector could produce a different estimate.

Conventional institutional sources do not dominate the entire citation pool. Among the 25 most-cited domains, academic and scientific publishers and platforms accounted for 14.6% of citations; news organizations had limited representation, while government domains accounted for 33%. Beyond that concentrated group lies the much larger long tail in which most domains appeared only once or twice. Separately, among the domains supplying the most URLs classified as AI-generated, the self-described AI research tool factually.co ranked eighth.

The exposure concentrates, but it isn’t contained. 28.0% of all unique source domains in the sample (1,754 of 6,258) had at least one source classified as AI-generated cited from them across the four engines.

Why the Footnote Used to Mean Something

A citation has always done quiet, structural work. It makes a claim traceable and borrows credibility from the source to which it points. Generative search engines inherited that visual grammar — the marker, the linked domain, the appearance of substantiation. But the citation alone does not disclose how the source was produced, whether qualified oversight shaped it, or whether its claims are sound. This audit exposes that gap. It does not reveal what checks, if any, occurred inside the engines before the source appeared.

The paper’s broader finding sharpens the point. These engines don’t draw evenly from a wide set of publishers. They repeatedly cite a concentrated group of domains while distributing most other citations across a long tail: 59.1% of cited domains appeared once, and 16.5% appeared twice. A page containing AI-generated text need only enter the retrievable web and be selected as a source. This audit shows that such pages are being selected; it does not establish what editorial review they received or reveal the internal rules that selected them.

The study did not test professional credentials, entity resolution, or whether an answer engine can distinguish an established practitioner from an imitation. Its relevance to this record is narrower — and still consequential. The engines repeatedly cited pages containing text Pangram classified as likely or highly likely AI-generated. Among the 200 such pages the researchers manually reviewed, none disclosed AI use. Visible citation therefore cannot, by itself, be treated as proof of authorship, editorial oversight, or authority.

Answer Engine Authority addresses a different layer of the problem: making a real professional’s identity, work, and corroboration machine-legible. This study does not show that such infrastructure fails. It shows why identity signals and citations cannot be treated as self-authenticating. The source still has to be evaluated.

Sources

Mowafak Allaham and Nicholas Diakopoulos (Northwestern University), Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources, arXiv:2605.23684, submitted May 22, 2026 (arXiv v1 preprint): https://arxiv.org/abs/2605.23684

A Link Is Not a Witness.

A footnote can show where an answer came from. It cannot, by itself, show who created the source, how much synthetic assistance shaped it, whether a qualified person reviewed it, or whether its claims are sound. That missing provenance is the next question for SIA — the Intelligence Officer,
briefed on every edition of this record the morning it releases.

Every edition, in order, from No. 001 · A new edition releases daily, 05:30 CT.

Open the Record
‹ Edition No. 058 Edition No. 060 ›