The people who do AI visibility work for a living just told the industry what they think of its tools. They rate the data itself 4.20 out of 5. Their willingness to fund a dedicated platform: 3.19 out of 5. Same survey, same respondents, a full point of difference.
Duane Forrester surveyed 163 AI search visibility practitioners over three weeks in July, with results published August 13 by Search Engine Journal. GEO stands for generative engine optimization, the practice of showing up inside AI answers. The survey's finding is now being called the GEO trust gap: practitioners want the data and do not trust the tools selling it. Only 44 percent think buying a tool in this category is worthwhile.
The complaints are specific, and they deserve to be taken seriously rather than defensively. Here they are in the practitioners' own order.
What 163 practitioners actually said
Trust, accuracy, and opaque methodology led the open-text responses at 24 percent. ROI and attribution to business value followed at 20 percent. Non-determinism and variance, meaning the same check returns different answers on different days, came in at 15 percent. Synthetic prompts versus real demand trailed at 11 percent.
The written comments explain the problem more plainly than the ratings do. Respondents called prompt-list tracking a self-fulfilling prophecy with no denominator, and noted that visibility scores come from invented prompt lists rather than observed query volume.
That complaint describes the core of the problem. If a tool scores you against a list of prompts the vendor made up, your score measures the vendor's list first and your brand second.
Each complaint points at a real mechanism
Opaque methodology. A score you cannot reproduce is an opinion. If the tool will not show you the exact queries it ran, you cannot re-run them, compare two months, or defend the number in a meeting. We wrote about why visibility tools disagree: two tools with different invented prompt lists produce different scores for reasons neither tool can detect. Practitioners live with this every week.
Non-determinism. AI engines do not return stable answers. The same question can produce a different citation set tomorrow. A single snapshot therefore carries almost no information. What matters is the trend across repeated, documented runs, and most tools do not publish their run history.
ROI and attribution. Digiday reported the same week that with monitoring tools now commonplace and 73 percent of surveyed CMOs invested in AI visibility monitoring, the hardest part to prove is that presence drives revenue. Boards are asking for the connection and the industry cannot yet draw it. When no one can show what the spend produced, buyers stop trusting the tools they are paying for.
Synthetic prompts. This is the no denominator complaint. Real customers ask questions in their own words, at volumes you can count. Invented prompts have no volume behind them, so a visibility percentage computed over them is a percentage of nothing measurable.
The shortlist problem makes it worse
A separate study, published the day after the survey, makes the complaints concrete. Suganthan Mohanadasan analyzed 60 ChatGPT conversations for Search Engine Journal and found that the model writes brand names into its own search queries before it retrieves anything, then probes the sites of the brands it named. He estimates being named is worth roughly 33 times more than being merely findable.
A prompt-list tool cannot see that stage at all. It can tell you that you did not appear in an answer. It cannot tell you whether you were named and dropped during verification, or never named because the model does not associate your brand with your category. Those are different problems with different fixes, and the trust gap grows every time a tool reports the symptom as if it were the diagnosis. We break the shortlist mechanism down in [ChatGPT Decides Who to Recommend Before It Searches](https://parceit.com/blog/chatgpt-decides-who-to-recommend-before-it-searches).
What decision-grade would have to look like
The IAB published Measuring Visibility in the AI Era on August 3, 2026, and drew a line that maps exactly onto this survey. Directional measurement hints something might be happening. Decision-grade measurement is reliable enough to act on. Practitioners rated the data valuable and the tools untrustworthy, which is the market saying, in numbers: give us decision-grade or stop billing us for directional.
Decision-grade AI visibility measurement has four requirements, and each one answers a complaint from the survey.
Documented queries. The exact questions, published, so anyone can re-run them. This answers opaque methodology directly. A methodology page that lists every check and every query is the bare minimum.
Real demand as the denominator. Questions drawn from what customers actually ask, in their phrasing, rather than a list a vendor invented. Query data exists. Search console logs exist. Using them is a choice.
Repeated runs, stored raw. Every time the tool checks your visibility, it should save the raw results exactly as they came back. Anyone can then check a score against the answers behind it. Saving every run also shows the pattern over time. AI answers change from day to day, so one run can come back unusually high or unusually low for reasons that have nothing to do with your site. Several runs across several weeks show which changes are real and which are ordinary day-to-day variation. This answers the non-determinism complaint: a single score from a single run cannot tell you that.
Stated limits. Measurement that says outright what it cannot do. No tool on the market can close the loop from visibility to revenue today. What the evidence supports, and what the Digiday piece confirms from the CMO side, is that presence measurement is commoditized and attribution is unsolved. Anyone who claims otherwise is selling directional data with decision-grade confidence.
Where Parceit stands
We built our audit around these requirements before the survey made them news. The methodology is documented on our site. The audit crawls your site with real HTTP requests and checks more than 30 factors, and the queries are reproducible. You can read the methodology before you give us anything, including your email.
We will also say what our free audit is not. It is a diagnosis, not a fix, and a single run is a snapshot. The same limits apply to us that apply to everyone, and we would rather state them than have you discover them.
The survey's 4.20 says the market believes in the data. The 3.19 says the market is waiting for someone worth trusting with it. That gap is the entire product roadmap for this category, written by its own customers.
Run a free audit at parceit.com. Read the methodology first. That order is the point.