Restricted analysis, made public daily.
Declassified under standing order Edition No. 067 Tuesday, September 8, 2026

Your Site Has Access Settings.

Have You Read What They Allow?

The open web now has a consent grammar for AI, a Google toggle that reached every website on August 31, 2026, and, beginning September 15, 2026, Cloudflare onboarding defaults that record the training answer for new ad-supported domains. Three layers of permission, and an open question of how fully any owner understands what each one records.

Cloudflare’s announced September 15, 2026 changes put a consequential question into website onboarding: which kinds of automated access should a domain permit? The company’s July 1, 2026 post states the plan: “For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.” Its August announcement says selecting the ad-supported option at onboarding sets Training to Disallow; non-publisher onboarding adds no blocks or disallows by default.

Cloudflare’s reasoning, published July 1, 2026, is candid about the bind it means to break. A small site owner, the company writes, faces a “Faustian bargain: either show up in search and let AI train on you, or risk losing discoverability.” And the post names who profits from the bind: “This unfairly advantages incumbent search providers if they use the same bots for both search and training” — an advantage, it adds, that incentivizes newer players to be evasive as they try to close the gap. The cure is a taxonomy: Search, Agent, and Training, managed separately. And the post is explicit about how far that enforcement grammar reaches for owners who choose the strictest setting: “Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training” — whether through the new controls or the legacy Block AI bots service.

Cloudflare’s August 21, 2026 announcement refines those mechanics, and softens them. Bot Preference Sync, announced that day with availability promised for the following week, writes the dashboard’s crawler preferences into the domain’s public robots.txt file. At onboarding, an ad-supported site can check a single box — “I monetize from pages with ads on this domain” — which sets Training to a recorded Disallow preference: in Cloudflare’s words, “you stay in search while keeping your content out of model training,” and cooperating mixed-use crawlers that take its transparency step “can still access your content for search indexing.” For everyone else the August post is explicit: “new customers will not have any blocks or disallows added by default when they onboard a domain.” August refines July’s approach: qualifying mixed-use crawlers can retain search access under Training Disallow; crawlers that fail the transparency requirements remain subject to blocking. Outside the ad-supported case, onboarding adds no default blocks or disallows. The accounts are dated seven weeks apart, and this record quotes both. Owners who want out of the new defaults, the July post had already promised, “can easily mark this in their Security settings” any time before September 15, 2026. The narrower question is the one this record exists to ask: mark it where, having learned about it how?

The default is only the newest floor of a consent architecture that has been assembling for a year. On September 24, 2025, Cloudflare published its Content Signals Policy: an addition to robots.txt defining three machine-readable signals, search, ai-input, and ai-train, so that “a website operator can then optionally express their preferences” for how content is used after it is accessed. At publication, Cloudflare counted its managed robots.txt feature already active “for over 3.8 million domains.”

Google built the layer above it. A Search Console control announced June 3, 2026, and rolled out to all websites worldwide as of August 31, 2026, per the post’s own footnote, lets website owners “decide if they want their site to appear in and help ground responses in our generative AI Search features (like AI Overviews, AI Mode or AI Overviews in Discover).” The scope is stated plainly on both sides of the choice: a site that opts out “will not receive traffic or impressions from our generative AI features,” and the setting “will not be used as a ranking signal” outside them. Consent, formalized — for the owner who knows the setting exists.

Sixty-six editions of this record have asked how the machines weigh what they read. This one asks who decided what they may read at all.

The floor, measured

An earlier peer-reviewed study documents awareness and compliance problems in crawler controls — Liu, Luo, Shan, Voelker, Zhao, and Savage, accepted to the ACM Internet Measurement Conference (IMC ’25), its measurements predating every announcement above. On the owner’s side of the file, the study surveyed 203 professional artists and found that “59% of the artists (119) had not heard about robots.txt prior to our study.” A majority had never encountered robots.txt, the file Cloudflare uses to publish its consent signals. On the machine’s side, the study’s measurements found compliance uneven in a specific pattern: seven major crawlers — Amazonbot, Applebot, CCBot, ClaudeBot, GPTBot, Meta-ExternalAgent, and OAI-SearchBot — respected the file, while one, Bytespider, “fetched the robots.txt file but did not respect it.” In the active test of AI assistant crawlers, ChatGPT’s and Meta’s built-in crawlers complied; “for the 23 third-party crawlers, most of them did not respect the robots.txt file,” and 20 of the 23 never fetched it at all. And of 1,875 top-10,000 sites behind Cloudflare, measured against October 2024 rankings, “only 107 (5.7%) sites enable Cloudflare’s Block AI Bots option.”

The bounds of each layer deserve exact statement. Cloudflare’s new defaults are written for newly onboarding ad-supported domains, though the July post’s most-restrictive-rules enforcement also changes how multi-purpose crawlers are treated for existing customers who had already selected a Training block. The Content Signals Policy expresses machine-readable preferences and includes a reservation-of-rights provision; it does not itself block access. Google’s control governs only appearance and grounding in its generative AI Search features. And the study sits beneath all three as measurement of robots.txt awareness and crawler compliance as they stood before these announcements; its own frame is training-data consent for content creators, and reading its third-party-crawler findings as a fact about the answer layer is this record’s observation, not the paper’s thesis. Three layers, one measurement, and no interchangeable scopes. What they share: a preference can be recorded without the owner fully understanding what it commits them to.

The clinical read — Decision Exclusion at the permission layer

The grammar does not enforce itself. The consent signals are real, machine-readable, and published at scale. In the study’s measurements, one major crawler fetched the robots.txt file and ignored it, and 20 of 23 third-party assistant crawlers never opened it. Those measurements demonstrate uneven compliance with robots.txt directives in the study’s tests. They do not measure compliance with the later Content Signals Policy.

The toggle requires owner awareness. A setting can only carry a choice for the owner who knows it exists. In the study’s survey, which predates these announcements, 59% of the 203 professional artists had never heard of robots.txt, the file the consent grammar rides on, and only 5.7% of measured Cloudflare sites had enabled its separate, older Block AI Bots option. Low adoption does not itself establish ignorance; it is the reason awareness cannot be assumed.

The preference takes shape at onboarding. Cloudflare’s August announcement says selecting the ad-supported option sets Training to Disallow. Qualifying mixed-use crawlers can retain search access; those that fail its transparency requirements remain subject to blocking. A control can record a preference without establishing how fully the owner understands its consequences. Digital Derangement Syndrome’s Decision Exclusion characteristic names the risk this creates, an entity absent from the set a system draws on, as a consequence to verify per system and access route, not an outcome these posts alone can diagnose. The same holds for Trust Transfer Failure, the contributing characteristic: a channel narrowed at Training is a channel to check, not a verdict.

Upstream of recognition

This is something quieter and earlier than the recognition failures this record usually files: the permission layer that recognition depends on, being written by defaults, grammars, and toggles. An enforced block can prevent a particular crawler from retrieving a page. Whether that restriction affects an answer depends on the system, the access route, and the other sources available to it. Permission therefore deserves deliberate review alongside entity architecture. That is why Answer Engine Authority treats the permission record as part of its first phase: before any question of how well an entity is described comes the question of which machines are permitted to read the description, and on what terms.

The consent architecture arriving this month — a grammar the reader may skip, a toggle whose setting deserves deliberate review, a default that files its answer at onboarding — makes the permission record a live surface of the web, one owners should inspect across their public signals and platform settings. A control can record a preference without establishing how fully the owner understands its consequences. The bargain Cloudflare named is real. The work this month leaves behind is reading your own file — and deciding, deliberately, what it should say.

Sources

Cloudflare, “Your site, your rules: new AI traffic options for all customers,” July 1, 2026 (the September 15, 2026 defaults for new onboarding domains, the Search/Agent/Training taxonomy, the Faustian-bargain framing, and the most-restrictive-rules enforcement) — blog.cloudflare.com.

Cloudflare, “Say it once: introducing Bot Preference Sync,” announced August 21, 2026, availability promised for the following week (dashboard preferences written to robots.txt; onboarding mechanics for ad-supported and non-publisher sites) — blog.cloudflare.com.

Cloudflare, “Giving users choice with Cloudflare’s new Content Signals Policy,” September 24, 2025 (the search / ai-input / ai-train signals and the 3.8-million-domain managed robots.txt figure) — blog.cloudflare.com.

Google, “New opportunities, control and insights for website owners,” June 3, 2026, updated August 31, 2026 (the Search Console control for generative AI Search features, its worldwide rollout, and its stated scope) — blog.google.

Enze Liu, Elisa Luo, Shawn Shan, Geoffrey M. Voelker, Ben Y. Zhao, and Stefan Savage, “Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers,” accepted to the ACM Internet Measurement Conference (IMC ’25) (the 59% survey figure, the crawler compliance measurements, and the 5.7% blocking figure, all as measured in the study) — arXiv:2411.15091.

Do you know which side of the toggle
your newest page is standing on right now?

September 15, 2026, is an onboarding line, not a launch date. Cloudflare’s August announcement says selecting the ad-supported option at onboarding sets Training to Disallow; non-publisher onboarding adds no blocks or disallows by default. Either way, the recorded answer operates whether or not its owner has read it. Answer Engine Authority begins one layer below content: entity architecture, and now the permission record itself,
held deliberately. If you cannot say which signals your own domain
is sending, that is a question for SIA — the Intelligence Officer,
briefed on every edition of this record the morning it releases.

Every edition, in order, from No. 001 · A new edition releases daily, 05:30 CT.

Open the Record
‹ Edition No. 066 Edition No. 068 ›