HomeInsights
Build

Should You Block or Allow AI Crawlers? The 2026 Decision Matrix

Block the wrong bot and you vanish from AI answers. Allow everything and you hand your content to models for free. Here is how to decide, crawler by crawler, in 2026.

Man in white shirt and patterned vest sitting in a black leather chair at a table, resting chin on fist.

Eden John

Founder, SkyScale

7 min read

Published

July 30, 2026

Updated

July 30, 2026

Decorative

What changed (30 July 2026): AI crawlers are no longer one undifferentiated swarm. Training bots, search bots and user-fetch bots now identify themselves separately, which means the old all-or-nothing choice is gone. You can block what takes and keep what gives, if you know which is which.

Table Of Content

Quick summary

The block-or-allow question has a wrong answer at both extremes. Blocking every AI crawler removes you from AI search; allowing every one feeds training models that rarely send anything back. The right move is granular: block the training crawlers that take without returning traffic, and keep the search crawlers that put you in front of buyers.

  • AI crawlers split into training, search and user-fetch bots
  • Training bots crawl heavily and refer almost no traffic
  • Search bots are how you appear in AI answers, so keep them
  • Blocking all AI bots is the costly mistake many made in 2024
  • Your decision depends on whether content is your product
Audience Icon

Who this is for

This is for owners, marketers and technical leads deciding what to put in their robots.txt, who want a reasoned answer rather than a blanket rule copied from a forum.

  • Marketers who need to stay visible in AI answers without giving everything away
  • Technical leads weighing content protection against AI search reach
Evidence base document icon

Evidence base

Across 200+ AI visibility audits we ran between October 2024 and June 2026, one of the most common self inflicted wounds was a robots.txt that blocked the very crawlers a business needed to be found by. Owners had reached for a blanket block to protect their content and quietly removed themselves from AI search in the process, usually without realising the two were different decisions.

Research methodology icon

Methodology

We reviewed how each major AI provider names and separates its crawlers, matched those against what independent traffic data shows each one actually does, and built the decision logic below from what consistently served businesses best given their goals. The aim is a framework you can apply to your own robots.txt, not a one size rule.

Limitations warning icon

Limitations

The crawler landscape changes fast, new bots appear and providers rename or split existing ones, so treat the specific names here as current examples rather than a permanent list. Robots.txt is also advisory, not enforcement, so the guidance below assumes well behaved crawlers and pairs with other controls where genuine protection is needed.

Office whiteboard displaying a 2026 decision matrix comparing which AI crawlers to allow or block for website access.

The dilemma: protect your content or stay visible?

The tension is real and worth stating plainly. On one side, letting AI crawlers in means your carefully produced content becomes training fuel for models that may then answer your buyers' questions without ever sending them to you. On the other, blocking those crawlers can wipe you out of the AI answers where a growing share of buyers now start, which is its own kind of loss.

Faced with that, plenty of businesses reached for the simplest lever in 2024 and blocked everything with an AI in its name. It felt safe. It was often a mistake, because it treated a nuanced set of very different bots as a single enemy. The good news for 2026 is that the tools to make a smarter choice now exist, and the choice is no longer binary.

The framing our AI search team uses with clients is that your robots.txt is a business decision wearing a technical costume. It is not really about bots; it is about which AI surfaces you want to be found in and what your content is worth to you, and once you have answered those two questions the technical rules mostly follow. Treating it as pure plumbing is exactly how the wrong crawler ends up blocked.

First, stop treating AI crawlers as one thing

The single most important shift is to stop thinking of AI crawlers as one swarm. At heart these are still web crawlers, the automated bots that the Wikipedia entry on web crawlers describes as systematically browsing and fetching pages, and which have always respected robots.txt as an advisory signal. What is new is that they now come in distinct types with very different purposes.

Search Engine Land's guide to AI crawlers lays out the categories cleanly. There are training crawlers that gather broad content to train models, such as GPTBot, ClaudeBot and Google-Extended. There are AI search crawlers that index pages so they can be cited in AI answers, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot. And there are user-fetch bots that grab a page when someone asks an assistant about it directly.

Each big assistant now runs this split. Appearing in ChatGPT means keeping OpenAI's search bot allowed even if you block its training bot, and the same logic governs staying visible in Perplexity, which leans heavily on live retrieval rather than training. The provider you most want to reach should shape which crawlers you prioritise.

That distinction is the whole game. Blocking a training crawler and blocking a search crawler are completely different decisions with completely different consequences, and lumping them together is what leads businesses astray. Once you see the categories, the block-or-allow question stops being one question and becomes several smaller, easier ones.

What training crawlers take, and how little they give back

Start with the training crawlers, because this is where the case for blocking is strongest. Their job is to hoover up content to train models, and the volume is staggering relative to what they return. Cloudflare's data on the crawl-to-click gap found that training now drives around 80% of AI bot activity, dwarfing search and user actions, and that the ratio of crawls to referred visitors is wildly lopsided.

The figures are hard to argue with. In that data, one major training crawler was fetching on the order of tens of thousands of pages for every single visitor it referred back, and even the more restrained ones took roughly a thousand pages per referral. In plain terms, training crawlers are a heavy withdrawal from your content with almost no deposit in return traffic.

A concrete example is CCBot, the crawler behind Common Crawl, whose open dataset feeds a great many AI models. Its own CCBot documentation shows how simple it is to disallow in robots.txt, and blocking it is a reasonable move if your worry is your work ending up in training corpora, since so many models draw on that one source. For a business whose content is a genuine competitive asset, blocking the pure training crawlers costs you little visibility and protects what matters.

What AI search crawlers give you, and why blocking them hurts

Now the other side. AI search crawlers are not taking your content to train on; they are indexing it so they can cite you when a buyer asks a question. Blocking these is how businesses accidentally delete themselves from AI search, and it is the expensive error to avoid.

This is where the granularity matters most. As Search Engine Journal reported when Anthropic split its Claude bots into separate crawlers, you can now disallow ClaudeBot for training while still allowing Claude-SearchBot for search visibility, and Anthropic warns directly that blocking the search bot may reduce your visibility and accuracy in results. OpenAI mirrors this with GPTBot, OAI-SearchBot and ChatGPT-User doing three different jobs.

The report made the wider point sharply: many sites that blocked all AI crawlers in 2024 inadvertently removed themselves from AI search citations, a costly mistake as that traffic grows. If being found in AI answers matters to you, and for most businesses chasing leads it does, the search crawlers are the ones you want to keep the door open for. Understanding how ChatGPT selects and cites its sources makes clear why locking out the search crawler is self defeating.

The key move: block training, keep search

Put those two halves together and the headline strategy for most businesses writes itself. Block the training crawlers that take your content without returning value, and keep the search crawlers that put you in front of buyers. It is not a compromise so much as the correct reading of two genuinely different situations. Keeping the door open is only half the job, though, because once a search crawler can reach you, your content still has to be structured so it can be lifted cleanly, which is where structured data for answer engine optimisation does the heavy lifting.

That default holds for the large middle of the market: service businesses, local operators, ecommerce brands and B2B companies whose goal is to be discovered and chosen. For them, AI search visibility is worth far more than the marginal protection of blocking a search bot, and the training bots offer nothing in return worth keeping. Making your content easy for the search crawlers to use is exactly what generative engine optimisation is built around.

The same logic holds provider by provider. If a meaningful share of your buyers use Claude, staying present in Claude's answers means allowing its search bot while you remain free to block its training bot, and the pattern repeats across every assistant that has separated the two. You are not making one decision so much as the same small decision several times, once per provider.

There is nuance underneath, of course, and the next blog in this series goes crawler by crawler. But the principle is stable enough to act on today, and it corrects the most common and most damaging error we see.

The 2026 decision matrix

The right call still depends on what you are optimising for, so run your situation through three cases.

If your business depends on being discovered and chosen, which covers most companies reading this, allow the AI search crawlers without hesitation and block the pure training crawlers if content protection matters to you. Your visibility in AI answers is the asset; protect it first. This is the default, and departing from it should be a deliberate choice, not an accident of a copied robots.txt file.

If your content is the product, such as a publisher, a research firm or a business whose written work is the thing people pay for, the balance tips further toward protection. You may reasonably block training crawlers across the board while still allowing search crawlers so buyers can find you, accepting that some models will still reference you second hand. Here the calculus weighs the value of your content as intellectual property against the reach you would lose, and reasonable owners land in different places.

If you are somewhere in between, decide page by page rather than site wide. It is entirely legitimate to let everything crawl your marketing and product pages while blocking crawlers from a proprietary knowledge base or premium library. Your robots.txt does not have to give one answer for the whole domain, and a serious AI visibility audit will show you which sections are earning citations worth keeping open and which are pure give-away.

How to actually implement your decision

Once you have decided, the mechanics are straightforward. You name each crawler with a user-agent line and allow or disallow it, and because the bots are now separated, you can be precise. Search Engine Land's guide to robots.txt for SEO in 2026 is a sound reference, and its core advice is to keep the file simple and avoid over-restricting, since a clumsy block does more harm than the threat it was meant to stop.

Two cautions matter. First, robots.txt is advisory: well behaved crawlers respect it, but it is a request, not a lock, so anything that genuinely must stay private needs real access controls rather than a polite note. Second, disallowing a crawler is not the same as keeping a page out of an index, so if your goal is exclusion rather than crawl control, pair your rules with the appropriate page-level directives. If robots.txt feels like unfamiliar territory, the fundamentals sit inside the broader complete guide to AI in SEO, which places crawler control alongside everything else that feeds your AI visibility.

Test after you change anything, because the cost of a mistake here is silent. A wrong line does not throw an error; it just quietly stops the crawler you wanted, and you find out weeks later when your citations fade. Auditing your own generative AI search visibility regularly is how you catch that early.

Common mistakes to avoid

The biggest mistake is the blanket block, disallowing every AI user-agent in one sweep and erasing yourself from AI search to stop training you barely slowed. The second is the opposite, allowing everything without thought and handing genuinely valuable proprietary content to training corpora for nothing.

The third is treating robots.txt as security, when it is advisory and cannot enforce anything against a crawler that ignores it. The fourth is setting and forgetting: the crawler landscape shifts, new bots appear, and a file written in 2024 may be fighting last year's battle. Review it as part of your regular technical checks, the same way you would any other part of your AI SEO foundation, and it will keep serving the decision you actually meant to make.

None of this is set and forget, and the businesses that get it right treat their crawler policy as a living decision rather than a one-time config. Our case study on a business that rebuilt its visibility began by untangling exactly this kind of robots.txt confusion, and if a blanket block has already cost you ground, the way back is the one we map in winning back visibility lost to AI search.

Implementation checklist

Use this list to audit and improve your AI visibility after reading this guide.

  • List every AI crawler currently hitting your site from your server logs
  • Sort them into training, search and user-fetch categories
  • Allow the AI search crawlers so you stay visible in AI answers
  • Block pure training crawlers only if content protection genuinely matters
  • Decide page by page for proprietary or premium sections
  • Keep the robots.txt file simple and avoid over-restricting
  • Pair blocks with page-level directives where real exclusion is needed
  • Re-test and review the file as part of every technical audit

Sources and references

Primary sources, official documentation, research and SkyScale audit data cited in this article. in this article.

Frequently Asked

Should I block AI crawlers or allow them?

Decorative

It depends on the crawler. Block the training crawlers that take your content to train models and send almost no traffic back, but allow the AI search crawlers that index your pages so you can be cited in AI answers. Blocking everything is usually a mistake, because it removes you from the AI search results most businesses want to appear in.

What is the difference between a training crawler and a search crawler?

Decorative

A training crawler, such as GPTBot or ClaudeBot, gathers content to train a model. A search crawler, such as OAI-SearchBot or Claude-SearchBot, indexes your pages so an assistant can cite them when answering a question. They are separate bots you can control separately, and they have opposite implications for your visibility.

Will blocking AI crawlers protect my content?

Decorative

Only partially, and only from well behaved bots. Robots.txt is advisory, so compliant crawlers will respect it but others may ignore it, and blocking a crawler does not remove already-trained content from a model. For genuine protection you need real access controls, not just a robots.txt rule.

Does blocking AI bots hurt my SEO or AI visibility?

Decorative

Blocking training bots generally does not hurt your visibility. Blocking search bots does, because those are the crawlers that let you appear in AI answers. Many businesses that blocked all AI crawlers in 2024 unintentionally removed themselves from AI search, which became costly as that traffic grew.

Can I block some AI crawlers and allow others?

Decorative

Yes, and that is the recommended approach. Because providers now name their training, search and user-fetch bots separately, you can disallow one while allowing another in the same robots.txt file. That granularity is what lets you protect your content from training while staying visible in AI search.

How often should I review my robots.txt for AI crawlers?

Decorative

Review it as part of your regular technical audits, at least a few times a year. The crawler landscape changes quickly as providers add, rename or split bots, so a file written a year ago may be blocking the wrong things. Re-testing after any change is essential, since a wrong rule fails silently.

Authorship and review

Man in white shirt and patterned vest sitting in a black leather chair at a table, resting chin on fist.

Written by

Eden John

· Founder, SkyScale

 LinkedIn profile

Eden leads SkyScale's Generative Engine Optimisation practice, focused on getting brands cited inside ChatGPT, Perplexity, Google AI Overviews and Gemini.

Relevant experience: Shipped 100+ AI visibility audits across B2B SaaS, professional services and ecommerce between Q4 2024 and Q1 2026, tracking citation patterns across the four major answer engines.

Credentials: Master of Business Administration (MBA) · Founder, SkyScale · 100+ AI visibility audits · GEO, AEO and AI SEO specialist

Smiling young man with curly dark hair in a maroon T-shirt crosses his arms indoors.

Reviewed by

Lachlan McDonald

· AI Search & Data Engineering Reviewer

 LinkedIn profile

Lachlan reviews SkyScale's AI search and data engineering content, focused on technical accuracy, methodology, retrieval logic, data quality and source-evaluation claims.

Relevant experience: 6 years of experience across AI search and data engineering, reviewing technical systems and source-selection claims for accuracy, reliability and methodological soundness.

Credentials: Master of Data Science · Bachelor of Software Engineering (Honours) · AI search and data engineering specialist

Last reviewed March 27, 2026
This is the block containing the Collection list that will be used to generate the "Previous" and "Next" content. You can hide this block if you want.
Ai visibility icon

AI Visibility
Report

3 business days. No credit card required, reviewed by a human.

Real Client Results

What you can expect to gain

+1,975%

more clicks from search

Benarrivati

£2,262

revenue from ChatGPT

Avenue Cookery

Google CTR lift

Vision One

+462%

more search impressions

SkyScale

See how we did it