Key Takeaways
- Bot Preference Sync generates or updates robots.txt from your AI bot config, prepending managed directives to keep existing rules intact.
- Category-wide controls let you allow or block Search/Agent bots while setting Training to disallow with transparency checks for mixed-use crawlers.
- Robots.txt remains voluntary, so pairing sync with enforcement like AI Crawl Control is essential to close the compliance gap.
Table of Contents
Cloudflare Makes AI Bot Preferences a Single Source of Truth
Cloudflare has removed the divergence between what a site says in robots.txt and what its enforcement layer actually blocks. Bot Preference Sync, now available across the Free-to-Enterprise plan stack, prepends AI bot preferences for Search, Agent, and Training to the live robots.txt file and can be switched off at any time.
The shift targets a persistent operational fault line: a robots.txt file can mark a crawler as disallowed while edge enforcement still fails to block that same crawler. When those layers disagree, some crawlers treat the mismatch as permission to disregard preferences or attempt to bypass enforced rules.
Your controls should reflect your strategy, which is why we’ve been building tools to give you visibility and choice at every layer.
Inside Bot Preference Sync: Search, Agent, and Training Controls
Bot Preference Sync closes the gap by generating or updating robots.txt from the AI bot configuration already set in the Cloudflare zone-level dashboard. For sites that already publish a robots.txt file, Cloudflare prepends the managed directives instead of replacing existing content, preserving any legacy Disallow rules.
Search, Agent, and Training Controls
- Search and Agent: Allow, block ad-served pages only, or block everywhere.
- Training: Disallow writes a no-training preference while preserving search access for crawlers that clear the transparency bar.
- Existing robots.txt: Managed rules are prepended, not replaced.
New customer sites have Bot Preference Sync enabled by default. Existing customers still using the legacy managed robots.txt feature are prompted to review and confirm their preferences before transition.
Publishers that depend on ads can select ‘I monetize from pages with ads on this domain’ during onboarding. That sets Training to Disallow by default while leaving the site available to search crawlers.
For everyone else, Cloudflare does not add any blocks or disallows unless the customer chooses them. Because the sync operates category-wide, it does not read individual custom rules with more complex logic.
Cloudflare maintains the bot lists through BotBase and its public bots directory. Customers with special one-off arrangements can turn off the sync and maintain a manually tailored file.
The transparency test for mixed-use crawlers
Mixed-use crawlers that blend search and training behind a single user agent now face a stricter verification requirement. To avoid being blocked when Disallow Training is set, they must honor the no-training preference, provide an opt-out from AI summaries, offer URL-level visibility into trained pages and search metrics, and show that disallowing training does not hurt traditional search.
Cloudflare publishes examples of crawlers that meet these conditions in the AI bot transparency section of Cloudflare Radar. Operators that do not provide transparency remain blocked when training is disallowed.
Voluntary robots.txt, Real Visibility Risk
Cloudflare’s technical documentation states the boundary without ambiguity: robots.txt expresses a site owner’s preference, but it does not by itself prevent access. Compliance depends on crawler operators, and some may disregard the directives entirely.
That is why Cloudflare positions robots.txt as one layer to be paired with AI Crawl Control for enforcement. Bot Preference Sync solves the consistency problem, not the compliance problem.
Adobe Experience League documents a related diagnostic called ‘Traffic Blocked by robots.txt’ that flags pages allowed for the wildcard user agent but disallowed for specific AI agents. It checks six major agents including ClaudeBot, GPTBot, OAI-SearchBot, OAI-User, PerplexityBot, and Perplexity-User.
The tool surfaces findings at line level and reports two numbers: total affected URLs and blocked agents. That level of detail matters because blocking a single high-traffic page can remove it from AI-generated responses, reducing citations and brand exposure.
Declared crawlers are only part of the traffic
Across the AI ecosystem, declared and identifiable crawlers are only the visible portion of data collection. Some operators rotate IPs or spoof user agents, so Bot Preference Sync works best as a policy control for cooperating crawlers rather than a substitute for active bot management.
For performance and SEO teams, this creates a direct feedback loop between crawler policy and AI-driven referral performance. A site that blocks Training without syncing robots.txt may think it is protected while still serving training crawlers; a site that blocks too broadly may vanish from AI answers entirely.
Crawler Policy Is Now a Performance Control
An inconsistent crawler policy used to be a hidden tax on visibility; now it is an explicit operational failure. For teams managing AI-driven crawler policies that need to scale, programmatic SEO AI automation is how Andres SEO Expert approaches the problem — talk to us about aligning your AI visibility strategy.
Frequently Asked Questions
What is Cloudflare Bot Preference Sync?
Bot Preference Sync is a Cloudflare feature that aligns a site’s robots.txt with its actual enforcement rules, preventing mismatches where a crawler is disallowed in robots.txt but not blocked at the edge. It generates or updates robots.txt based on the AI bot configuration in the Cloudflare dashboard, with controls for search, agent, and training crawlers.
How does Bot Preference Sync handle an existing robots.txt file?
Cloudflare prepends the managed AI bot directives to the existing robots.txt file rather than replacing it. This preserves any legacy Disallow rules you have already set, ensuring your custom preferences remain intact while gaining the new AI-focused controls.
What is the difference between Search, Agent, and Training controls?
Search and Agent controls let you allow a bot, block only ad-served pages, or block it everywhere. Training controls, when set to Disallow, communicate a no-training preference while still allowing search access for crawlers that meet transparency requirements. This distinction helps you manage AI crawling without harming your search visibility.
What does the transparency test require for mixed-use crawlers?
Mixed-use crawlers that combine search and training under one user agent must honor no-training preferences, provide an opt-out from AI summaries, offer URL-level visibility into trained pages and search metrics, and prove that disallowing training does not hurt traditional search to remain unblocked when Training is disallowed.
Does robots.txt block crawlers by itself?
No. robots.txt only expresses site owner preferences and compliance depends on crawler operators. It does not enforce access control on its own, so Cloudflare recommends pairing it with enforcement tools like AI Crawl Control to ensure that disallowed bots are actually blocked.
Why is it important to sync robots.txt with edge enforcement?
When robots.txt and enforcement disagree, some crawlers ignore the preferences or attempt to bypass enforced rules. Syncing them closes this gap, giving you consistent, accurate control over AI bots and reducing the risk of losing visibility in AI-generated responses due to mismatched rules.
What are the limitations of Bot Preference Sync for undisclosed crawlers?
It works as a policy control for cooperating crawlers, but some AI operators rotate IPs or spoof user agents, making their bots invisible to declared crawler lists. For more robust protection, you need active bot management to catch these hidden crawlers.
