Say it once: Introducing Bot Preference Sync
Cloudflare's new Bot Preference Sync automatically aligns your robots.txt file with your AI bot policies for Search, Agent, and Training. Easily manage which bots access your content without maintaining static files.
Source: Cloudflare Blog · August 22, 2026 at 2:01 PM · AI-assisted report
Single-sourceKUALA LUMPUR, 22 AUGUST 2026 —
Listen to this article
DomainFork Audio · read aloud
Cloudflare Launches Bot Preference Sync to Align Robots.txt with AI Bot Policies
Market Impact
Cloudflare announced on 21 August 2026 the rollout of Bot Preference Sync, a new feature that automatically updates a website’s robots.txt file to reflect the AI bot preferences set in the Cloudflare dashboard. The service is available to all customers, from the Free tier to Enterprise, and is designed to simplify the management of search, agent, and training bot traffic.
The feature builds on Cloudflare’s earlier July 1, 2026 launch that introduced managed robots.txt values and edge‑enforced blocks for AI training crawlers. Bot Preference Sync extends that capability by synchronising the AI bot configuration with the robots.txt file in real time. When a site owner enables the sync, Cloudflare prepends the appropriate directives to the existing robots.txt, preserving any custom Disallow rules already in place.
Background: The Need for Unified Bot Controls
Web operators face a spectrum of bot‑traffic objectives. Some prioritize discoverability, allowing search engines to index all content, while others seek to protect intellectual property by blocking training bots. The coexistence of multiple mitigation layers—robots.txt directives, Cloudflare’s Bot Management rules, and third‑party bot lists—has historically led to inconsistencies. For example, a site might disallow a crawler in robots.txt but fail to block it at the edge, prompting compliant crawlers to ignore the directive or malicious ones to bypass it.
Cloudflare’s earlier solution addressed two layers simultaneously: a managed robots.txt value that disallowed major training crawlers and edge‑enforced blocks for those same crawlers. Bot Preference Sync now unifies all three AI bot categories—Search, Agent, and Training—into a single, coherent policy reflected both in the dashboard and in the robots.txt file.
How Bot Preference Sync Works
When a site owner sets a preference for a bot category, Bot Preference Sync writes the corresponding directive to robots.txt. For Search and Agent, the options remain: Allow, Block on pages that serve ads, or Block everywhere. For Training, the new Disallow option writes a “Disallow: /” directive for training bots, while still permitting cooperating mixed‑use crawlers that provide transparency to access the site for search indexing.
The system pulls from Cloudflare’s BotBase, a continuously updated list of verified bots. When a category is blocked or disallowed, the relevant bots are added to the robots.txt file. The list is refreshed periodically to account for new bots or changes in bot behaviour. Verified bots are publicly listed in the AI bot transparency section of Cloudflare Radar, allowing site owners to see which bots have met the transparency requirements.
Impact on Malaysian and Regional Web Operators
For Malaysian e‑commerce sites, the ability to allow all content to be crawled and trained on can improve product visibility in AI‑powered shopping assistants. Conversely, publishers and media outlets that rely on ad revenue may prefer to keep their articles out of training datasets while maintaining search visibility. Bot Preference Sync gives operators the flexibility to tailor their policies without maintaining multiple static files.
Regional operators can also benefit from the transparency framework. Bots that provide additional information about their identity and data usage are granted access even when “Disallow Training” is set. This encourages responsible bot operators and can reduce the risk of unauthorized content use. The feature’s default activation for new customers further simplifies adoption, while existing customers using legacy managed robots.txt will be prompted to transition.
Stakeholder Perspectives
Cloudflare spokespersons emphasized that the feature “ties together the preference you set with the preference you publish.” They noted that the sync is optional; operators with complex, custom security rules can disable it and manage their robots.txt manually. The company also highlighted that publishers or ad‑supported sites can now have a different default policy from other site owners, reflecting the varied needs across the industry.
Data and Figures
- Bot Preference Sync is available to all Cloudflare customers, from Free tier to Enterprise. - The feature was launched on 1 July 2026 as part of a broader AI bot management strategy. - The Disallow option for Training writes a “Disallow: /” directive to robots.txt. - BotBase is updated periodically to reflect changes in bot lists. - Verified bots are listed publicly in Cloudflare Radar’s AI bot transparency section.
Forward‑Looking Outlook
Cloudflare plans to continue refining its bot‑management ecosystem. The company is exploring additional transparency metrics and will likely expand the range of bot categories in future releases. For Malaysian operators, the immediate benefit is a streamlined, automated approach to managing AI bot traffic, reducing administrative overhead and aligning on‑site policies with external bot behaviour. As AI‑driven search and content generation become more pervasive, tools like Bot Preference Sync will play a crucial role in balancing discoverability, revenue, and content protection.
Related: Sea