Should I Block AI Crawlers in robots.txt?

Short answer

It depends on whether your content is the product. A publisher has a real interest in not having work ingested for free. A local business that wants to be recommended has close to the opposite interest, because a blocked site is harder for an assistant to describe.

Run the robots.txt Tester

Read your robots.txt the way a crawler does.

The crawlers people mean

Several distinct bots get lumped together as AI crawlers, and they do different jobs, which matters if you are deciding selectively.

User agentRun byBroadly what for
GPTBotOpenAITraining data collection
OAI-SearchBotOpenAISurfacing results in ChatGPT search
ClaudeBotAnthropicTraining data collection
PerplexityBotPerplexityAnswering with citations
Google-ExtendedGoogleGemini training, separate from Search
CCBotCommon CrawlAn open dataset many models train on

The genuine case for blocking

It is not paranoia, and dismissing it would be dishonest.

If your content is your product, being ingested means a model can answer the question your content answers without anyone visiting you. News, reference material, tutorials, recipes, research: all have a direct interest in not being repackaged.

There is also a straightforward fairness argument. Your work becomes commercially valuable to a company that neither asked nor paid, and blocking is the only lever most site owners have.

Google-Extended is a particularly clean decision, because it governs Gemini training separately from Search. You can block it without affecting how you rank.

The cost most people do not connect

Blocking has a consequence that is easy to miss if you think of these crawlers only as scrapers.

When someone asks an assistant to recommend a plumber in their town, or which software does a particular job, the assistant answers from what it has read. A site that blocked every AI crawler has removed itself from that consideration.

For a business whose website is marketing rather than product, that is a real loss and it is growing. The visitor who would have found you through an assistant simply does not.

This is why the right answer differs by business rather than being universal. A local service business and a magazine have genuinely opposed interests here.

Selective blocking, and honest limits

You can split the decision rather than treating it as all or nothing. Blocking the training crawlers while allowing the ones that answer with citations is a coherent position: no free training data, but stay findable.

Two limits worth knowing. First, robots.txt is advisory. Well-behaved crawlers honour it and anything scraping in bad faith ignores it, so blocking is a request rather than a control.

Second, blocking now does not remove what was already collected. Models trained on earlier crawls retain it.

And be sceptical of anything sold as making you more visible to AI. Publishing a file that describes your business is not a substitute for other people writing about you, which is what models actually learn from. Google has said it does not use llms.txt for its AI features, and Venbit does not generate such files.

Common questions

Does blocking AI crawlers hurt my SEO?+

Not your Google Search ranking. Google-Extended governs Gemini training separately from Search, so blocking it does not affect how you rank. What it does affect is whether assistants can read your site to recommend you.

Which AI crawlers should I block?+

It depends on your interests. If your content is your product, the training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended) are the ones to consider. If you want to be recommended, allowing the ones that answer with citations makes sense.

Does blocking remove my content from models already trained on it?+

No. Blocking prevents future collection. Anything gathered in earlier crawls remains in models already trained on it.

Will an llms.txt file help me appear in AI answers?+

Be sceptical. Google has said it does not use llms.txt for its AI features, and publishing a file about yourself is not a substitute for other people writing about you. Venbit does not generate these files.

Is robots.txt enforceable against AI crawlers?+

No, it is advisory. Reputable crawlers honour it; anything scraping in bad faith ignores it entirely. It is a request, not access control.

Sources

A sitemap helps machines find your pages

It does nothing for the visitor who lands on one and cannot find what they came for. A Venbit agent reads your pages and answers them directly, by text or by voice. Free to start, no card.