← Back to blog
·12 min read·Heidi Macomber

The 40% GEO Gain Is a Myth. Here Is What Actually Predicts AI Visibility.

A new meta-study of 45 GEO papers finds the famous "40% visibility gain" is a lab artifact, not a business outcome. Here is what the evidence actually says about ranking in AI search, and why core retrieval signals beat prompt-tracking every time.

GEOAI SearchAI VisibilityStructured DataMeasurement

If you have been to a marketing conference, read an SEO blog, or sat through a vendor pitch in the last year, you have heard the number. "GEO delivers a 40% gain in AI search visibility." It is the headline stat of the Generative Engine Optimization industry. Agencies quote it. Tools advertise it. LinkedIn influencers build entire carousel posts around it.

It is not real.

A new meta-analysis published on arXiv (paper 2607.14035, Olivier Martinez, July 2026) reviewed 45 studies on Generative Engine Optimization and found that the famous 40% figure traces to a single metric in a single test configuration, and that the academic literature places it in its lowest confidence tier. The study's broader finding is more damning: while rewriting content can change how an already-retrieved page gets cited, almost none of the studies prove the page gets retrieved at all, and nearly none measure whether the changes produce durable clicks or conversions.

In other words, the GEO industry has been measuring the wrong thing, inflating the results, and selling you the difference.

This article breaks down what the evidence actually says about AI search visibility, why the metrics most tools sell you are measuring noise, and what core signals reliably determine whether your business gets cited by ChatGPT, Perplexity, and Google AI Overviews.

What the meta-study actually found

The paper, titled "A Critical Review of Generative Engine Optimization Benchmarks," synthesizes 45 published studies on GEO. Its findings fall into three categories.

The 40% gain is a lab artifact

The headline 40% visibility gain that appears across the GEO marketing literature originates from a single metric called Position-Adjusted Word Count (PAWC), measured in one specific test configuration across a limited set of queries. When the meta-analysis applied standard confidence grading, PAWC-based gains landed in the lowest tier. Other metrics in the same studies showed smaller, less consistent effects.

The number is not fabricated. It was measured. But it was measured in a narrow lab setting, on a metric that does not correspond to any business outcome, and it has been repeated so often that it has taken on the appearance of an established fact.

Citation changes do not prove retrieval

Several GEO studies show that rewriting content (adding more authoritative language, restructuring paragraphs, adding citations within the text) can change how a page is cited once it has been retrieved by an AI engine. But the meta-study points out a fundamental gap: almost none of the studies prove the page gets retrieved at all.

This is like measuring whether a better storefront display increases purchases after ruling out whether anyone walked down your street. Citation optimization matters, but only if the retrieval pipeline can find your page in the first place.

Business outcomes are almost never measured

Of the 45 studies reviewed, nearly none tracked clicks, conversions, lead generation, revenue, or any durable business metric. The vast majority measured citation frequency or position within an AI-generated answer, then stopped. There is no evidence in the academic literature that GEO interventions produce lasting traffic or revenue gains.

Why prompt-tracking is the wrong instrument

A separate article published in Search Engine Journal (July 19, 2026) reached a compatible conclusion from a different direction. It argued that the dominant method for measuring AI search visibility, typing prompts into ChatGPT, Perplexity, and Google AI Overviews and counting brand mentions, is fundamentally flawed for three reasons:

  • An AI prompt is not a keyword. Keyword research assumes stable intent behind a search term. AI prompts are conversational, context-dependent, and infinitely variable. A prompt list you invent in a conference room does not match what real users actually ask.
  • Mention counts do not correlate with business outcomes. Being mentioned in an AI answer does not mean the user clicks through, contacts you, or buys. In fact, AI answers are designed to satisfy the user without clicking. High mention counts with zero clicks is not a win. It is a vanity metric.
  • Prompt-tracking is not reproducible. AI engines personalize answers based on user history, location, and session context. Two people typing the same prompt can get different answers. Running the same prompt twice can produce different results. You cannot build a measurement framework on a non-reproducible signal.

The article calls for tracking three things instead: presence (whether you appear at all), recommendation share (how often you are the recommended answer versus a competitor), and brand accuracy (whether what the AI says about you is correct).

The Google credibility gap

Complicating the measurement problem, Google itself is making claims about AI search performance without providing the data to back them up.

On July 18, 2026, Google's Nick Fox stated that AI features in Search send "billions of clicks weekly" to websites. The claim was widely repeated in marketing media. But Google provided no baseline, no denominator, no methodology, and has not published the underlying data.

"Billions of clicks weekly" sounds impressive until you ask: out of how many total queries? What percentage of AI answers include a clickable source? How many of those clicks convert? Without a denominator, the number is not a metric. It is a press release.

This is not a Google-specific problem. Every AI search engine has an incentive to publish numbers that make their features look good for publishers, because publisher cooperation (via crawl access and structured data) is what makes the features work. Independent, site-level measurement is the only way to know what is actually happening with your business specifically.

What actually predicts AI visibility

If the 40% GEO gain is a myth, prompt-tracking is flawed, and platform claims are unverifiable, what should you actually measure?

The answer, backed by the retrieval pipeline architecture of every major AI search engine, is foundational signals. These are the inputs that determine whether an AI engine can find, parse, and understand your content at all. They are measurable, reproducible, and they correlate directly with retrieval.

Signal 1: Crawlability (Can the AI engine reach your page?)

Every AI search engine starts with a crawl. If your robots.txt blocks AI crawlers, or fails to allow them, your site does not exist in that engine's index. Period.

In July 2026, Google renamed NotebookLM to Gemini Notebook and changed the crawler user-agent. Sites that hardcoded the old user-agent in robots.txt will stop working in August 2026. This is not a hypothetical. It is a calendar event. If your robots.txt references the old user-agent, your crawlability is degrading this month.

Check for: robots.txt allow/disallow rules for GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and the current Gemini Notebook user-agent. XML sitemap submission. Server response codes.

Signal 2: Entity data (Can the AI engine understand who you are?)

When an AI engine retrieves your page, it looks for structured data (JSON-LD schema) to understand what entity the page represents. Without schema markup, the AI has to guess. With it, there is no ambiguity.

The minimum viable entity data for any business is:

  • Organization or LocalBusiness schema with your legal name, URL, logo, and verified social profiles
  • NAP (Name, Address, Phone) that is identical across your website, Google Business Profile, and all directories
  • Same-as links connecting your website to your social profiles and other authoritative references

If your NAP is inconsistent across the web, AI engines may not recognize your listings as referring to the same entity. This is one of the root causes of the 64% local business error rate.

Signal 3: Content structure (Can the AI engine parse your content?)

AI engines extract information from pages. They look for semantic HTML (proper heading hierarchy, nav, main, footer), question-and-answer formatting (FAQPage schema), and clear topical structure.

Pages with H1 tags, hierarchical H2/H3 headings, and FAQPage schema get cited more often because the AI's language model can extract clean, structured answers from them. Pages that are visual-only (large images, minimal text, JavaScript-rendered content) are functionally invisible.

Signal 4: Content accuracy (Is what the AI says about you correct?)

This is the newest signal, and it is the one the GEO studies completely ignored.

A July 2026 vendor test found that AI tools returned at least one false fact about 64% of UK high street retailers. The errors were not in the content the businesses published. They were in how the AI engines synthesized information from multiple sources, filled gaps with guesses, and presented the result as fact.

You cannot control what an AI engine generates. But you can audit what it is generating right now, and you can strengthen the underlying signals that reduce the likelihood of errors. Accurate, consistent entity data across your site and the broader web gives the AI fewer opportunities to get it wrong.

The Germany ruling: why this matters more than ever

On July 14, 2026, Germany's media regulator (ZAK) classified Google AI Overviews and Perplexity as "content providers," not neutral search engines. This strips their liability shield under the EU Digital Services Act, because AI-generated responses count as the providers' own editorial content. A Munich court separately held Google liable for false claims in an AI-generated answer.

This is not a European regulatory footnote. It is the beginning of a legal framework where AI answers are treated as published content, subject to accuracy standards and liability. When that framework reaches US courts, and it will, businesses that have audited their AI visibility will have documentation of what the engines were saying and when. Businesses that have not will be starting from zero.

The combination is stark: AI engines get business information wrong 64% of the time, they are now legally accountable for those errors in growing jurisdictions, and the metrics most vendors sell you to track this are either inflated (the 40% gain), flawed (prompt-tracking), or unverifiable (platform claims).

How Parceit measures AI visibility

Parceit was built to measure the core signals that actually determine AI visibility. Our audit does not track prompts or count mentions. It crawls your website and scores the four signal categories described above:

  • Structural (30 points): HTML fundamentals, semantic markup, heading hierarchy, meta tags
  • Entity (25 points): JSON-LD schema, organization data, NAP consistency, same-as links
  • Content (20 points): FAQPage schema, article structure, question/answer formatting
  • Crawlability (25 points): robots.txt, sitemap, AI bot access, server response codes

Each category maps to a specific stage in the AI retrieval pipeline. If your site scores low on crawlability, AI engines cannot reach your content. If it scores low on entity data, they cannot identify you. If it scores low on content structure, they cannot extract clean answers. If it scores low on all three, they guess, and the 64% error rate tells you how that goes.

The audit runs in under 60 seconds. There is no signup wall. You get your score, a breakdown of every signal we checked, and a prioritized list of what to fix first.

Run your free audit at parceit.com.

Frequently asked questions

Is GEO (Generative Engine Optimization) real?

GEO is a real concept but the marketing claims around it are not supported by evidence. The academic literature shows that content optimizations can affect citation behavior but do not reliably produce retrieval or business outcomes. The underlying signals that GEO tries to optimize are the same signals traditional SEO has measured for years: crawlability, structured data, and content quality. There is no separate GEO discipline. There is only SEO done completely.

Should I track AI mentions by typing prompts into ChatGPT?

Prompt-tracking has limited value as a spot check but is not a reliable measurement framework. AI prompts are not keywords. Mention counts do not correlate with business outcomes. Results are not reproducible across users or sessions. Track presence, recommendation share, and brand accuracy instead, and measure the base signals that determine whether you can be retrieved at all.

What is the 40% GEO gain and why is it misleading?

The 40% visibility gain cited across GEO marketing traces to a single metric (Position-Adjusted Word Count) in one test configuration, placed in the lowest confidence tier by the 2026 meta-analysis of 45 GEO studies. The gain was measured in a lab, on a metric that does not correspond to clicks, conversions, or revenue. It is not a business outcome and should not be used to evaluate AI visibility investments.

How do I know what AI engines are saying about my business?

Run a Parceit audit to check your core signals, then manually test a handful of real customer queries in ChatGPT, Perplexity, and Google AI Overviews. Focus on whether the information is accurate (right address, hours, services) rather than whether you are mentioned. Accuracy problems trace back to entity data gaps that the audit will identify.

Want to know how your site scores?

PARCEIT's structural audit engine crawls your website and checks all of these signals in under 5 seconds. Find out exactly what AI search engines see.

Run your free audit