The clearest answer: use separate rules for training, user-directed access, and search

The best AI crawler policy examples begin by separating three activities that publishers often incorrectly treat as one: training crawlers that collect material for model development, input crawlers that fetch pages in response to a user or tool, and ordinary search crawlers that discover and index content. Google’s robots.txt specification supports the use of user-agent groups such as ai-train, ai-input, and search, with values of yes or no for each. This is more precise than blocking every unfamiliar bot, although support for these group names is not automatic: a named crawler still has to honor the directives, and publishers should test behavior rather than assuming a declaration has been applied.

Also worth reading: How Should Publishers Control AI Crawler Permissions in 2026? · How Do Modern Publishers Create a Defensible AI Publishing Policy Template in 2026? · What AI Policy Should Publishers Adopt Before AI Enters Their Book Workflow?

A strong policy normally starts from a business decision rather than a technical reflex. Publishers who depend on search discovery may permit conventional search indexing while restricting model training or retrieval. Sites licensing content may permit selected AI partners. Sites concerned about attribution, bandwidth, or server load may use controls at the CDN or WAF level. In short, there is no universally correct example: the useful example is the one that maps crawler permissions to the publisher’s content strategy, legal obligations, and tolerance for automated access.

robots.txt: a simple, widely recognized policy model

A conventional robots.txt policy is the easiest starting point for a small publisher because it is plain text, maintained alongside the website, and recognized by a mature class of crawlers. It may create a named group for AI training crawlers and disallow their access, while another group permits search indexing. It may also include a wildcard rule for unlisted agents, but a wildcard has a major limitation: it can accidentally block beneficial services, including some search engines or tools not considered when the file was written.

The major weakness is enforcement. robots.txt is chiefly a request to voluntarily respected automated clients; it is not an authentication system and does not itself prevent a determined client from requesting a public page. Directives do not create a takedown mechanism, and their interpretation varies when bots claim different identities. Therefore, a file containing “User-agent: ai-train” and “Disallow: /” is only meaningful if the crawler actually recognizes that user-agent token and honors the rule.

For important assets, robots.txt should be paired with HTTP access controls, rate management, or application controls. This is especially important for files that contain noindex directives but should not be downloaded indiscriminately, because robots.txt controls crawling rather than indexing. Conversely, disallowing a URL does not necessarily remove an already indexed URL from results. The best examples are simple, testable, and reviewed on a schedule rather than copied uncritically from another publisher.

Featurerobots.txt approachCDN or WAF policyAccess-controlled content
Setup effortUsually lowModerateHigher
Enforcement strengthVoluntaryStrong technical controlStrongest
Suitable forBasic crawler guidanceTraffic filtering and bot managementLicensed, private, or premium material
Typical costFreeOften included, then paid tiersUsually paid or contractual
Main weaknessBots may ignore itMay classify legitimate users as botsCan exclude users and reduce sharing
## Cloudflare-style controls: stronger enforcement with more operational tradeoffs

Cloudflare’s newer crawler controls illustrate the second major policy category: managing automated traffic at the network edge rather than relying only on robots.txt. Its announcement described additional options for AI traffic, including controls for blocking selected AI crawlers, allowing them, or limiting how they may interact with a site. Cloudflare has also discussed default blocking and commercial terms such as pay per crawl, showing why crawler governance is becoming part of hosting economics rather than merely a website preference.

This model offers stronger enforcement because traffic can be challenged, delayed, blocked, or subjected to commercial terms before content is served. It can address abuses that robots.txt cannot, such as high request rates, expensive pages, or automated extraction that ignores voluntary restrictions. However, the operational consequence is greater complexity. Rules can be bot-specific, user-agent-specific, path-specific, or account-specific, and maintainers need to understand whether a legitimate user’s request includes an AI agent.

A CDN policy is not automatically “better for publishers.” It costs more to configure and monitor, can interfere with search and AI accessibility, and may vary by plan, region, and Cloudflare product release. Search Engine Journal has also examined concern that AI crawler rules could block Googlebot, illustrating how an apparently precise rule can produce unintended search effects. The appropriate example is therefore a documented allow or deny matrix, tested against named agents and at least one normal browser request, with a rollback procedure.

Content signals and metadata: useful additions, not substitutes

Content-Signal and related standards attempt to express machine-readable preferences about uses such as ai-train, ai-input, and search. The value of these signals is that they can communicate intent directly instead of requiring an operator to infer a publisher’s wishes from one generic disallow rule. They can also provide a standardized field for systems that choose to honor machine preferences, and they may be useful to future content-management and licensing platforms.

The limitation is equally important: a metadata signal has value only when downstream services discover, interpret, and respect it. It cannot protect content that has already been downloaded, establish an enforceable license, or automatically compensate a publisher. Different crawlers may use different file locations and schemas, while content-management systems may strip unknown metadata. Publishers should therefore verify whether the signal is emitted, whether it can be fetched without crawling prohibited resources, and whether known partners actually act on it.

Metadata works best as one layer in a policy stack. A publisher might expose content preferences in metadata, repeat selected restrictions in robots.txt, and enforce network-level limits for serious abuse. It should also record the date of each change because crawler support evolves quickly. The November 2024 and March 2026 research discussions around public-data rules demonstrate active debate, but industry debate should not be confused with consensus or universal technical compliance.

A practical publisher policy: build, test, document, and revisit

The first practical step is to inventory agents that already request the site. Server logs, CDN analytics, and security dashboards can identify user agents, IP ranges, request frequency, paths, and response codes. A reasonable initial review should examine at least the prior 30 to 90 days, because short samples may miss seasonal crawlers. The publisher then classifies every important automated client as search, training, retrieval, monitoring, security, accessibility, affiliate activity, or unknown. This exercise often reveals that “AI crawler” is not a single category.

The second step is to choose default behavior. For many content publishers, a sensible starting position is to permit ordinary search and selected user-facing tools, restrict bulk training where use is not authorized, and investigate unknown agents rather than either allowing or blocking them blindly. Publishers licensing material should avoid applying a broad denial that conflicts with active agreements. The wording should distinguish crawling from indexing and user-directed retrieval from corpus acquisition.

Testing should include fetching robots.txt, examining page delivery, reviewing logs, and checking whether known user-agent tokens are being filtered. Policies should be versioned with dates, owners, approved agent names, and an explanation for each exception. A quarterly review is a useful minimum, while a site undergoing rapid AI-policy changes may review monthly. There is no universal legal safe harbor in robots.txt, so legal review remains important when the site is the official source, processes sensitive data, or operates under a specific contract.

Common mistakes and the reasons they backfire

The most common mistake is copying another publisher’s file without checking whether the business models match. A policy designed for a subscription archive may be unnecessarily restrictive for a public-interest database, while a policy intended to stop model training may interfere with search indexing. Another error is assuming that disallowing a page in robots.txt removes it from search results; indexing and crawler access are related but distinct systems.

A second mistake is relying on the User-Agent header without verification. Automated clients can misidentify themselves, legitimate tools may rotate infrastructure, and a token alone may not identify the service operating it. Maintainers should compare documented user agents and verified IP ranges where the provider supplies them, but they should not assume every matching IP is trustworthy. Rapidly changing infrastructure makes static allowlists fragile.

The third mistake is treating restrictive policy as zero-cost. Blocked traffic can reduce bandwidth and server expense, but it can also remove referral opportunities, prevent citations, break accessibility or product integrations, and complicate monitoring. A rule that blocks all AI access may disappoint users who want the publisher’s material included in an answer, while a rule that permits everything may expose content to extraction without attribution or payment. Publishers should measure referrals, search visibility, crawl volume, server costs, and complaints before and after implementation.

Finally, administrators often fail to separate robots.txt from authentication. Private files should use access control; premium content should use entitlements; and public pages may use crawler policy. Replacing a subscription check with “Disallow” weakens the original business model rather than protecting it.

Alternatives, timing, and cost expectations

Publishers have several alternatives. An open policy maximizes discoverability but gives up detailed control over reuse. A robots.txt-only policy is inexpensive and transparent but weak against noncompliant clients. A CDN rule offers stronger technical control, often at no additional charge on some plans, although advanced controls and managed services may require paid configuration or subscriptions. A licensing program may produce revenue, but it introduces contracts, usage reporting, attribution, and enforcement work. Removing content from public access is strongest technically but conflicts with sharing, archive, and public-interest goals when applied indiscriminately.

Small sites can usually begin at zero direct cost by using robots.txt, logs, and documented tests. Larger publishers should budget staff time for inventory, testing, monitoring, and legal review, even if the software is free. CDN products can range from included features to paid enterprise arrangements; any exact price must be checked with the provider because plans and crawler controls change. Payment per crawl is still an evolving commercial concept, not a stable category with one market price.

Action should be considered when AI requests become material—for example, when automated traffic exceeds ordinary monitoring thresholds, consumes disproportionate bandwidth, produces repeated unlicensed extraction, affects search performance, or conflicts with an active content agreement. Immediate action is appropriate after a confirmed abuse incident or a rights complaint. Waiting may be reasonable for a low-traffic experimental crawler whose behavior is documented and whose access benefits the publisher. The policy should be revisited at least quarterly and after major crawler changes, with a documented review date rather than permanent assumptions.

The best example is an auditable decision, not the strictest file

There is no single best AI crawler policy because crawler behavior, publisher economics, and legal exposure differ by site. The strongest general example combines named search and AI groups in robots.txt, machine-readable preferences where supported, CDN enforcement for important paths, and access controls for private content. It permits useful discovery while distinguishing model training from user-directed retrieval, then tests whether actual traffic follows the declared policy.

The central test is whether a publisher can explain, in plain language, who may crawl the site, for which purpose, under what limits, and what happens when a rule is violated. If the answer is unclear, the policy is incomplete regardless of how sophisticated its file looks. Publishers should retain evidence such as test dates, observed status codes, crawler IP documentation, approved exceptions, and the responsible owner. That record becomes more valuable as AI agents proliferate and standards remain unsettled.

For an AI Publishing Consultant, this is a governance problem first and a software problem second. Technology can enforce preferences, but it cannot decide whether training, search, retrieval, monitoring, and commercial reuse are acceptable to a particular publisher. A measured policy that permits some responsible access can produce better relationships than a blanket ban, while stronger enforcement protects costly resources. The goal is not maximal restriction; it is predictable, measurable control aligned with the publisher’s actual goals.