← Back to blog
·8 min read·Heidi Macomber

What Blocking Google's AI Crawlers Actually Does (And What It Doesn't)

Publishers are blocking Google-Extended thinking they've opted out of AI Overviews. They haven't. Here is what each crawler actually controls, what blocking really does, and how to decide whether your site should opt out.

Robots.txtGoogle-ExtendedAI SearchCrawlabilityAI VisibilityGoogle AI OverviewsPublishers
Share:

In August 2026, People Inc. CEO Neil Vogel told investors that blocking Google's crawlers was "100% on the table" but that the company was "clearly not turning this off now." The reason was straightforward: 21% of People Inc.'s traffic still comes from Google search, down from 25% last quarter. If they block Google's AI crawler, they lose search referrals too. So they are stuck.

People Inc. is not alone. USA Today, Reuters, CNN, the New York Times, BBC, and Yahoo have all blocked Google-Extended, the robots.txt token that controls whether Google can use crawled content for AI training. Several, including USA Today and Reuters, are publicly considering blocking Googlebot entirely, which would remove them from Google Search results completely.

But here is what most publishers and brands get wrong: blocking Google-Extended does not opt you out of AI Overviews. It never did. And the confusion is causing real damage, because brands are making decisions based on assumptions about crawler control that do not match how Google's systems actually work.

This guide breaks down what each Google and AI crawler actually does, what blocking each one actually controls, and how to make an informed decision instead of a panicked one.

The four Google crawlers you need to understand

Google operates multiple crawlers, and they serve different purposes. Treating them as one thing is the root of most crawler-blocking mistakes.

Googlebot

This is Google's primary web crawler. It discovers pages, reads content, and builds the index that powers Google Search. If you block Googlebot, your site disappears from Google Search entirely. No organic results, no Top Stories, no Discover. This is the nuclear option, and almost no publisher has done it because the traffic loss is immediate and severe.

Google-Extended

Google-Extended is a separate robots.txt token that controls whether content Google has already crawled can be used for two things: training future Gemini models and grounding responses in Gemini Apps and Vertex AI. It does not control whether your content appears in AI Overviews or AI Mode. Google's own documentation states that Google-Extended does not affect a site's inclusion in Search and is not a ranking signal.

This is the critical misunderstanding. Blocking Google-Extended blocks AI training. It does not block AI Overviews. Publishers who blocked Google-Extended thinking they had pulled out of Google's AI answers are still showing up in them.

Googlebot-News

This crawler specifically feeds Google News. Blocking it removes your content from Google News without affecting regular search results. For news publishers, this is a surgical option that limits Google's use of content in news-specific surfaces without sacrificing broader search visibility.

APIs-Google

This crawler handles Google's API-driven products, including certain ads and feed-based services. It is separate from the main search index and rarely the focus of blocking decisions.

The new Search Console opt-out: what it actually does

In June 2026, under pressure from the UK's Competition and Markets Authority, Google began rolling out a new Search Console setting called the Search Generative AI control. This setting lets website owners exclude their content from AI Overviews, AI Mode, and Discover's generative AI features without leaving regular Search.

For the first time, this creates a genuine choice. You can stay in Google Search but remove yourself from Google's AI-generated answers. That is something robots.txt could never do before.

But there is a catch. According to NewzDash CEO John Shehata, nearly 1 in 6 US trending news queries (15.5%) and 17.46% of UK trending news queries now place Top Stories carousels inside AI Overviews. If you opt out of AI features, you may also be opting out of Top Stories placements that appear within AI Overviews. Google has not officially confirmed how the control handles traditional features embedded inside AI surfaces.

In other words, the opt-out is not as clean as it sounds. You might gain control over your AI presence but lose visibility in features you never thought of as AI.

Non-Google AI crawlers: GPTBot, CCBot, PerplexityBot

Google is not the only company crawling your site for AI purposes. OpenAI's GPTBot, Common Crawl's CCBot, and Perplexity's PerplexityBot each have their own robots.txt tokens. Blocking or allowing each one is a separate decision.

Some publishers have taken a blanket approach, blocking all AI crawlers. Others, particularly content creators and brands, are doing the opposite: allowing all AI crawlers and actively optimizing for AI visibility. This divergence between publishers (retreating from AI) and brands (leaning in) is one of the defining splits in search right now.

Who is blocking what

According to data compiled by Exploding Topics in August 2026, here is where major publishers stand:

  • BBC, Yahoo, New York Times, CNN: Blocking Google-Extended (AI training only), still in Google Search
  • USA Today, Reuters: Blocking Google-Extended, publicly considering blocking Googlebot entirely
  • People Inc, Politico: Not yet blocking Google-Extended, but considering full crawler blocks
  • Nearly 80% of the top 1,000 websites block at least one AI crawler

The pattern is clear: most major publishers have blocked AI training but are hesitating on the bigger step of blocking search. They are caught in what industry observers are calling a prisoner's dilemma. No single publisher can afford to be the first to cut Google off completely, because whoever does loses search traffic while competitors keep theirs.

Why brands should think differently from publishers

Publishers sell content. Their business model is based on getting traffic to pages with ads. When AI summaries answer questions without sending users to the source, publishers lose pageviews and ad revenue. Blocking crawlers is a rational, if risky, negotiation tactic.

Brands are in a different position. A brand's goal is not pageviews. It is visibility, trust, and being top of mind when a potential customer asks an AI assistant for a recommendation. If your brand is invisible in AI answers because you blocked crawlers, you do not lose ad revenue. You lose customers.

The brands winning in AI search right now are the ones doing the opposite of publishers. They are allowing all crawlers, making their content easy to extract, and measuring whether they actually appear in AI answers. They are leaning in while publishers pull back.

How to actually decide

The decision tree is simpler than the industry debate makes it sound.

If you are a publisher whose revenue depends on pageviews

Blocking Google-Extended is reasonable. It prevents Google from training models on your content without sacrificing search traffic. But understand that it does not remove you from AI Overviews. If you want out of AI Overviews specifically, Google is rolling out a Search Console setting called Search generative AI that lets you opt out of AI Overviews, AI Mode, and Discover without leaving regular Search. The rollout is gradual: it launched in June 2026 for UK sites and is slowly expanding to US properties through the summer. Many site owners will not see this control yet, and for most brands it is a non-issue right now since the goal is to be visible in AI answers, not to hide from them.

If you are a brand whose goal is visibility and leads

Do not block any AI crawlers. Allow Googlebot, Google-Extended, GPTBot, CCBot, and PerplexityBot. Make your content easy to crawl and extract. The brands cited in AI answers are the ones whose content was available to crawl, structured for extraction, and specific enough to quote. Blocking crawlers to make a point about AI fairness is a luxury brands cannot afford.

If you are not sure where you stand

Run a visibility audit. Check whether your brand actually appears in AI answers across ChatGPT, Perplexity, and Google AI Overviews. If you are invisible, blocking crawlers will not fix that. It will make it worse. If you are visible and want more control over how you are portrayed, that is a content and structure problem, not a crawler problem.

The real risk: deciding without data

The biggest danger right now is not blocking too much or too little. It is making crawler decisions without knowing what they actually do to your visibility. People Inc.'s CFO, Timothy Quinn, noted that ad rates have gone up significantly even as traffic declined, because quality content is commanding a premium. That is a data-driven insight. Most brands making crawler decisions today do not have equivalent data for AI visibility.

Google's new generative AI performance report in Search Console shows impressions from AI features but carries no click data and no query-level detail. Like the opt-out control itself, this reporting is rolling out gradually and may not yet be available for your property. The CMA has required Google to provide more, including click-throughs and click-through rates, but that data is not available yet. Until it is, brands need third-party measurement to understand their AI visibility.

That is what Parceit does. A Parceit audit checks whether your brand appears in AI answers across every major AI search engine, measures how prominently you are featured when you do appear, and flags whether the information is accurate. Before you touch your robots.txt or Search Console settings, know what you are actually gaining or losing. Run a free audit at parceit.com.

Share:

Want to know how your site scores?

PARCEIT's structural audit engine crawls your website and checks all of these signals in under 5 seconds. Find out exactly what AI search engines see.

Run your free audit