← Back to blog

Redson Dev brief · PRIMARY SOURCE

ARTICLE#Dev#AI

Have it both ways: stay discoverable in search while disallowing AI training

Cloudflare Blog · September 15, 2026

Web publishers can now navigate the complex waters of AI-driven content consumption without sacrificing search visibility, ensuring their content is seen while its use remains controlled. This piece from Cloudflare introduces new mechanisms, including specific crawler controls and an "Accountable" designation, designed to establish a consistent framework with major tech platforms for distinguishing between general search crawling and AI model training. The core offering allows site owners to explicitly permit content indexing for search engines like Google and Microsoft, while simultaneously blocking the same content from being scraped for the training datasets of generative AI models. For a freelance designer in Austin, Texas, specializing in brand identity, this means her online portfolio can continue ranking high for "boutique logo design" without her unique visual creations being absorbed into an AI generator that might dilute her distinctive style. A small e-commerce shop in Portland, Oregon, selling artisanal soaps, can ensure its product descriptions and customer reviews drive organic traffic for "natural skincare," yet prevent that copy from being ingested by an AI model that could then generate identical-sounding product text for competitors. Similarly, an indie SaaS founder in Boston, Massachusetts, building a niche project management tool, can protect their detailed knowledge base and blog tutorials from being used to train a competing AI service, while still appearing prominently in search results for specific technical queries. These controls provide a critical layer of intellectual property defense in an increasingly AI-driven digital landscape. The practical implication is a crucial distinction between content discovery and content exploitation. Readers can implement these controls to manage how their digital assets interact with the evolving web. This empowers them to maintain their online presence and reach, which is vital for business and visibility, while asserting agency over their proprietary information. It's a pragmatic step towards a more nuanced approach to content rights in the age of large language models. To capitalize on this, consider visiting your site's `robots.txt` file or your CDN settings this week. Look for options or patterns that allow you to distinguish between search engine crawlers and AI training bots, specifically exploring directives that permit `Googlebot` or `Bingbot` while restricting `GPTBot` or similar AI agents. If your current provider offers specific configurations for AI crawlers, experiment with implementing these to shield your content from unsolicited training while preserving search discoverability.

Source / further reading

Learn more at Cloudflare Blog