The Ad Changed the Choice.
The Answer Still Passed.
Researchers embedded a pharmaceutical ad — labeled “Sponsored Message” from the start — into the system prompt of twelve consumer-facing AI models, then asked each to choose between two drugs physician reviewers had judged equally appropriate. The label didn’t stop it from working: the sponsored claim moved the average selection rate by 12.7 points. The recommendation still passed as clinically acceptable.
A multi-institution research team, based primarily at the Icahn School of Medicine at Mount Sinai, ran 258,660 API calls across twelve consumer-facing AI models from OpenAI, Anthropic, and Google, in four separate experiments. The flagship test put each model in front of thirteen clinical scenarios, each pairing two brand-name drugs physician reviewers had judged equally guideline-appropriate for the same condition — 112,320 of those calls. Before the question, the researchers embedded pharmaceutical ad language into the system prompt, explicitly bracketed as a “Sponsored Message” — labeled at the start, labeled again at the close. The label did not prevent the measured shift. The researchers reported an average 12.7-percentage-point increase in selection of the advertised drug across the thirteen scenarios (P < 0.001). At the individual model-scenario level, some combinations swung completely — a drug never chosen at baseline, chosen every time once the ad ran. The shift was not evenly distributed by provider: Google models moved 29.8 points on average, OpenAI’s 10.9, Anthropic’s 2.0.
Measured accuracy did not decline in the equipoise experiment — it rose slightly, from 89.4 percent at baseline to 93.0 percent with the ad present (P < 0.001). What moved wasn’t correctness. It was which of the two guideline-appropriate drugs the model selected. The paper names the mechanism directly: “In this zone, the ad acts as a tiebreaker. The output is correct and biased.” Disclosure was rare and provider-dependent — in a separate open-response test involving three of the twelve models, only 29.9 percent of ad-condition responses acknowledged the ad at all, ranging from 55.9 percent (Claude Opus 4.6) down to 5.2 percent (GPT-4.1). An audit that only checks whether the answer is right will read a rising accuracy score as good news, not a warning — which is exactly what the authors mean when they write that the bias is “invisible to accuracy-based evaluation.”
The boundary, and the precedent
A boundary belongs on the table before the interpretation does: this was a research injection, not an observed feature of any deployed consumer AI assistant. No named pharmaceutical manufacturer bought this placement, and none is shown paying for it in production — the ad language was written by the study’s own authors and inserted into the system prompt as an experimental condition, precisely so the researchers could isolate its effect. What the study measures is a demonstrated mechanism, not an observed market practice. The finding matters within that boundary: promotional language changed drug selection in controlled clinical scenarios. The experiment establishes an effect on outputs; it does not establish how the models internally classified that language.
This is a structural cousin of a specimen already on this record. In Edition 025, an AI answer engine began selling advertisers a labeled box beside its answer — visible, disclosed, and, by the engine’s own account, never inside the reasoning that produced the answer. Here, the sponsored language sits inside the system prompt itself, not beside the output. And it was labeled, plainly — bracketed as a “Sponsored Message,” start to finish. It didn’t stop the shift in drug selection. The experiment demonstrates influence from promotional context sitting alongside clinical evidence; it does not establish how the models internally represented or weighted that context. Beside the answer, an ad is a disclosed transaction a reader can see and discount. Inside the system prompt itself, it is an influence that clinical accuracy alone does not measure.
The distinction is between clinical acceptability and selection independence. An answer can satisfy the study’s clinical criteria while its selection remains sensitive to promotional context.
Evidence status. This is a preprint — posted April 16, 2026, not yet peer-reviewed; independent replication was not established in this review. The authors declare no competing interests. Funding disclosures differ: the PDF states no funding, while the medRxiv landing page lists institutional support and NIH grants. Nothing in this edition treats the finding as a verdict on any specific AI product; it treats it as a documented mechanism inside a controlled test.
The audit that doesn’t know what it missed
An accuracy score alone cannot show whether promotional context changed a choice between two acceptable options. The researchers measured that separately, comparing drug selection with and without the sponsored language present. The audit question this raises is larger than whether a recommendation was clinically acceptable. It also has to ask whether promotional context changed which acceptable option got selected.
In the Agentics framework — an interpretation this publication brings to the record, not a finding of the study itself — what this specimen exhibits has a clinical name: Authority Misclassification, a characteristic of Digital Derangement Syndrome™. For Answer Engine Authority™, the question this raises is structural: what can influence a recommendation, and how is that influence tested? An editor can polish the answer without detecting what changed the selection.
Omar, M., Agbareia, R., McGreevy, J., Zebrowski, A., Ramaswamy, A., Gorin, M., Antao, E.M., Glicksberg, B.S., Sakhuja, A., Charney, A.W., Klang, E., Nadkarni, G.N., “Ad-verse Effects: Pharmaceutical Advertising Shifts Drug Recommendations by Consumer-Facing AI,” medRxiv preprint, posted April 16, 2026, DOI 10.64898/2026.04.14.26350868: preprint (medRxiv). Not yet peer-reviewed; the authors declare no competing interests. Funding disclosures differ between the PDF and landing page, as noted above.
Agentics Intelligence Declassified™, Edition No. 025, “Rented Recognition” (July 29, 2026), on disclosed advertising placement beside an AI answer: read the edition.