← All writing

Before you block the bots, decide what your website is for.

A bot-blocking checkbox is a terrible place to outsource your business judgment. Decide what each page should do, then choose who gets access.

A button that says “block AI bots” lets you take a firm position without the inconvenience of knowing what it is.

What are you protecting? The research people pay to read? The service page you paid someone to help customers find? A customer document that should never have been public?

Those pages have different jobs. Giving them the same access policy because they share a domain saves you a decision. It can also sabotage the reason you put them online.

On September 30, 2026, Cloudflare announced its Pay Per Use beta. Participating AI companies offer payment for specified uses of content; publishers choose which offers to accept. It adds another option to a discussion often reduced to letting crawlers in or keeping them out.

Before choosing an option, write down what the page is supposed to accomplish. You can survive ten minutes without having an AI strategy.

Three pages, three different jobs

Consider a fictional consultancy with a service page, a free troubleshooting guide, and a paid research report. This is a decision exercise, not a client story.

The service page should help a suitable customer understand the offer and make an enquiry. You paid to explain what the company does. Making that explanation harder to find deserves a reason. An assistant accurately describing the service could help, even if the customer arrives later through a different route. An assistant inventing a guarantee could create a problem. Neither outcome is measured by counting crawler requests.

The free guide helps readers solve a small problem and decide whether they trust the company. Some readers will get what they need without contacting anyone. The internet offered this deeply inconsiderate feature long before AI. The question is how much reuse supports the guide’s purpose, and where the company wants to draw a boundary.

The paid report is itself a product. If another service supplies enough of its substance to replace a purchase, the company has a different concern. It might publish a useful abstract for discovery while keeping the full report behind authenticated access. It might consider a specific licensing offer. It does not have to apply the service page’s policy to the report.

Start with the outcome for each page. “More traffic” is incomplete. More qualified enquiries, successful self-service, and paid readership are different results. A policy that helps one can obstruct another.

The checkbox is doing a lot of editorial work

Fetching a page, including it in search, using it to answer a question, and using it for model training are different activities. The available controls do not always separate them neatly.

For example, OpenAI documents independent settings for OAI-SearchBot and GPTBot: the former supports search discovery, while the latter crawls content that may be used for model training. It separately describes ChatGPT-User visits initiated by users, for which robots.txt rules may not apply. That distinction belongs to this provider’s documented behavior; it is not a universal map of every AI service.

Ask whoever manages the website to translate each proposed switch into a specific effect: which crawler, which pages, which use, and what happens when a visitor requests the page through an assistant. “It blocks AI” is the label you already read. Ask for the explanation. Record what the provider does not let you separate.

Allowing access creates an opportunity for discovery. It does not buy a citation, a visit, or a customer. Blocking access also needs a business reason. A vendor’s urgent-looking banner is evidence that the vendor wants you to click something.

Your PDF cannot be secured by a strongly worded request

Crawler preferences, search visibility, and private access need different controls.

Google’s robots.txt guidance explains that a disallowed URL can still appear in search results. A robots.txt rule is also no place to put your faith in the confidentiality of a document. Private material needs access controls that refuse unauthorized requests.

For a public page you want excluded from Google Search, Google supports a noindex instruction. But the crawler must be able to retrieve the page to see it. Blocking the crawl while expecting Google to read the instruction is a configuration argument you will lose. Google’s noindex documentation spells out that dependency.

For the fictional consultancy, this means checking that the report’s download actually requires authorization. If someone can open the supposedly restricted PDF without signing in, the access policy has failed. The robots.txt file can be beautifully formatted while this happens.

Use the control that enforces the intended boundary. Then test the affected page and its downloadable files, not just the checkbox.

A monetization tab is not a revenue stream

Cloudflare’s announcement distinguishes Pay Per Crawl, which charges for access, from Pay Per Use, which pays for agreed downstream uses. Under the latter, buyers define their offers and report usage. Cloudflare says reporting is required by the program terms and that it checks whether reported uses map to enrolled publishers. That is not independent observation of every use.

For a content business, an offer may deserve examination. Read what counts as a paid use, what reuse the terms permit, how reporting works, and how participation ends. Check eligibility and actual offers before putting anticipated revenue in a forecast. “There is now a button for this” is a shit revenue assumption.

For a consultancy whose writing supports enquiries, the calculation may be different. Would the arrangement help the writing reach useful readers? Would charging for access get in the way? Who will administer it? The service page may earn its keep by helping someone hire you. It does not need a second career selling admission to itself.

You can decline an unattractive offer without deciding that all automated discovery is worthless. You can welcome discovery without accepting every proposed use of the work.

Write the policy for three real URLs

Choose a page that attracts customers, a page that distributes useful information, and a page or file whose access needs restriction. If the third category does not exist, leave it out. Do not invent a paid-content strategy to finish a worksheet.

Complete these fields for each:

Decision What to record
Purpose Who needs this page, and what useful outcome should follow?
Access Public, paid, or private? Include linked files and alternate URLs.
Discovery Which search or answer services should be able to find it?
Reuse Which uses are acceptable, unacceptable, or still undecided?
Controls Named provider settings, crawler preferences, and any enforced access restrictions.
Limits What those controls cannot distinguish or guarantee.
Evidence What you will observe to judge whether the policy helps.
Review Owner, review date, and the condition that would trigger an earlier change.

For the consultancy’s service page, a reasonable starting proposal is to permit supported search discovery, make a separate decision about training, and review qualified enquiries and obvious misrepresentations. For its paid report, the proposal might be a discoverable abstract, authenticated full access, and no licensing participation until an actual offer has been evaluated.

Have the person responsible for each page’s business outcome approve its row. They should be able to explain the tradeoff without forwarding a screenshot of the settings panel.

Check whether the choice helped

Before changing settings, record the current policy and enough baseline information to make a later comparison useful. Start with the logs and enquiry records you already have. Buy another analytics subscription when you can name the question those records cannot answer.

After the change, check that intended visitors can still reach the public pages, restricted files remain restricted, and observed crawler responses match the chosen rules. A request that merely claims a bot’s name does not establish its identity; use the provider’s documented verification method where available.

Choose a review period suited to the site’s volume. Look at useful enquiries, reader feedback, reported licensing revenue where applicable, and access problems. Keep observations separate from explanations: fewer visits after a policy change does not, by itself, prove the policy caused the decline. A quiet site may need more time before the numbers say anything worthwhile.

Leave the exercise with a decision for each page, a person responsible for it, and a reason to revisit it. If you still cannot explain what blocking a crawler would protect or cost, leave the damn switch alone until you can.

Product details and linked documentation checked October 1, 2026. Recheck provider controls and program terms before changing a live site’s policy.