You ran your brand through three AI visibility tools. You got three different scores. One says you are dominating. One says you are invisible. The third says something in between.
Which one is right?
Digiday reported this week that marketers are losing patience with expensive AI visibility tools that deliver inconsistent results. The frustration is understandable. You are paying for clarity and getting contradiction.
But the inconsistency is not a bug. It is architecture. Here is why tools disagree, why most of them cannot help it, and what the difference is between a number that is interesting and a number you can actually act on.
Why tools give you different scores
There is no single index behind AI search. Each engine, ChatGPT, Perplexity, Google AI Overviews, Gemini, Claude, builds its own retrieval pipeline, trains on different data, weighs sources differently, and applies different reasoning before generating an answer. Kevin Indig's AI Halftime Report, published this month via Search Engine Journal, found that 91% of AI citations appear in only one engine. The same query across four engines produces four different sets of cited sources.
That fragmentation is the root of the problem. Every AI visibility tool has to make choices about how to measure across engines that do not agree with each other. Those choices produce different scores for the same brand. Here is where the disagreement comes from.
Different query construction. Every tool has to decide what questions to ask the AI engines. Do they use your brand name alone? Your brand name plus a product category? Competitor comparison queries? Long-tail questions your customers actually type? Two tools using different query sets will get different answers, because the AI engines respond differently to different phrasings. If one tool asks "best CRM software" and another asks "is [your brand] a good CRM," the results are not comparable.
Different engines tracked. Some tools track ChatGPT only. Others add Perplexity and Google AI Overviews. A few include Gemini and Claude. If Tool A tracks two engines and Tool B tracks five, their scores measure different things. A brand that is strong in ChatGPT but invisible in Perplexity will look great on Tool A and mediocre on Tool B.
Non-deterministic AI responses. AI search engines are not databases. You do not query them and get the same result every time. The same question asked on Monday and Wednesday can return different sources, different rankings, and different brands cited. Tools that run queries once and report the result are capturing a single point in time. Run the same query again tomorrow and the answer may change. This is not the tool's fault, but it means single-run scores are snapshots, not measurements.
Different scoring methodologies. Some tools count raw mentions. Others weight by position within the answer. Others factor in sentiment or whether the mention is positive, negative, or neutral. Some calculate share of voice against competitors. Some blend all engines into a single number. Some report per-engine. Two tools with different scoring formulas will produce different numbers from the same underlying data.
Black-box proprietary scores. Several competitors have launched proprietary scoring metrics with proprietary names. These scores are not comparable across tools because the formulas are not disclosed. A "Brand Consideration Score" of 72 from one tool and a "Visibility Index" of 45 from another tell you nothing relative to each other. You cannot benchmark, you cannot reproduce, and you cannot verify.
The IAB drew a line: directional vs decision-grade
The IAB published "Measuring Visibility in the AI Era" on August 3, 2026, as part of Project Eidos. The framework introduces four dimensions of AI visibility: Presence, Prominence, Portrayal, and Persuasion. But it also introduces a quality standard that directly addresses the inconsistency problem.
The IAB distinguishes between two levels of measurement quality.
Directional data tells you something might be happening. Your brand was mentioned in some AI answers. Your share of voice went up or down. This is useful for spotting trends, but it is not reliable enough to make decisions on. If you cannot explain how the queries were constructed, if you cannot reproduce the results, and if you cannot compare scores across runs, you have directional data.
Decision-grade data is reliable enough to act on. It uses documented, consistent query construction. It is reproducible: running the same audit twice produces the same result, or the differences are explained and bounded. It is transparent: you can see exactly what was measured, how, and against which engines. The IAB explicitly calls out query construction as a disclosure requirement. If your tool cannot tell you how it formed its queries, your results are not reproducible.
This is the line that matters. Most AI visibility tools on the market today produce directional data. They count mentions, generate scores, and produce dashboards. But they do not disclose their query construction, they do not guarantee reproducibility, and they do not report per-engine results in a way you can verify.
That is why you get different answers from different tools. They are all producing directional data using different hidden methodologies, and none of them are held to a standard that makes their numbers comparable.
What you should actually be measuring
The IAB's four dimensions give you a framework for cutting through the noise. Instead of chasing a single blended score from a black-box tool, measure each dimension with a method you understand.
Presence: Can AI crawlers reach your site at all? This is a structural question, not a query question. Check your robots.txt. Verify that GPTBot, PerplexityBot, ClaudeBot, and Google-Extended are allowed. This is binary and verifiable. No AI visibility tool should disagree on this if it is actually checking your crawlability.
Prominence: When an AI engine retrieves candidate pages for a query, does it pick yours? This depends on your ranking strength and topical authority. You can check this manually by asking real questions in ChatGPT, Perplexity, and Google AI Overviews and seeing whether your site appears as a cited source.
Portrayal: When the AI does cite you, is what it says accurate? Does it describe your products, pricing, and features correctly, or does it hallucinate? A German court ruled in June 2026 that Google is directly liable for false claims in its AI Overviews, treating them as Google's "own words." Portrayal accuracy is now a legal concern, not just a marketing one.
Persuasion: When the AI mentions your brand, does anyone click through? Use GA4 referral data to see which AI engines send you traffic. The IAB reports that AI referral traffic is growing rapidly and consistently outperforms average engagement metrics. But referral data only shows clicks. It does not show citations where a competitor won and you lost silently.
How to evaluate an AI visibility tool
If you are evaluating tools, ask three questions.
First, can the tool explain how it constructs its queries? If the answer is no or "it is proprietary," you have directional data. That can be useful for trend-spotting, but do not make budget decisions on it.
Second, does the tool report per-engine results or a single blended score? A blended score hides the fragmentation. You need to know where you are strong and where you are absent. ChatGPT and Perplexity cite different sources 91% of the time. A single number averages away the most important information.
Third, is the methodology reproducible? If you ran the same audit next week, would you get comparable results? If the tool's results swing wildly between runs and the tool cannot explain why, you are measuring noise.
The bottom line
The reason your tools disagree is not that one is right and the others are wrong. It is that they are all measuring different things using different hidden methods, and calling the results by the same name.
The fix is not to find the one tool that is "correct." The fix is to demand decision-grade measurement: documented query construction, per-engine reporting, and reproducible methodology. That is what the IAB standard calls for. That is what makes a number worth acting on.
If your AI visibility tool cannot tell you how it built its queries, which engines it tracked, and whether the results are reproducible, you do not have a measurement. You have an opinion with a dashboard.
Run a free audit at parceit.com. The methodology is transparent, the queries are documented, and the results are reproducible. You see exactly what was measured and why, across every major AI search engine.