The popularity of the AI search platforms (Google AI mode, ChatGPT etc) has led to a mini boom in companies claiming they can track an organisation’s “AI visibility”. However, all is not what it seems. In my experience, in their sales pitches and marketing materials, the AI tracking companies are not challenging their less-informed customers’ assumptions and are glossing over the limitations of their tools – which is very close to being misleading.
Below is an outline of some of the key points I highlight when discussing AI tracking with our clients.
Why the “% of total conversations” number is not available
Underneath the branding, almost every AI visibility platform works the same way. They take a set of prompts, either seeded by you or expanded automatically from a topic, and fire them at the AI models on a schedule. They then read the responses and record whether your brand was mentioned, whether your site was cited, and how you compare with competitors. The results are rolled up into a visibility score or a share of voice figure.
In other words, the tools invent the questions and measure the answers. The demand side, the questions real people ask, is manufactured by the tool. Only the supply side, what the models happen to say back, is genuinely observed. The prompt list is a synthetic panel, not a picture of real usage.
So when your boss says: “we want to state that our brand is mentioned in a given percentage of AI responses about our topic”, that claim is a fraction. On top, the number of AI answers about the topic that mention your brand. On the bottom, the total number of AI answers about the topic. To state the percentage as a fact about the world, you need that bottom number, the denominator. And the denominator is the thing nobody outside the model owners can see.
Why can’t the total denominator be seen? Here are four reasons:
- The answers happen in private, one person to one model. There is no public index of AI conversations to crawl, the way web pages or search results can be gathered. Only OpenAI, Google and Anthropic see their own traffic.
- The questions are effectively unlimited. A topic is not one prompt but an unbounded set of ways to phrase it, each asked some unknown number of times. To weight the percentage properly you would need the real volume behind every phrasing, which is exactly the data the providers hold and do not release.
- The traffic is scattered and no one holds all of it. Prompts flow through the consumer apps, the APIs, Copilot, and dozens of built-on products, across several competing models. Even OpenAI cannot see Google’s traffic. A figure describing “all AI responses” has no single owner who could calculate it, even with full internal access.
- The answers are not fixed. This is “generative AI!”. The same prompt returns different responses across runs, model versions, accounts and locations. There is no stable, repeating result to count in the way a Google ranking can be counted.
This means that no tool can produce the true denominator, and no brand can honestly claim to know all AI responses.
The output still needs to be cleaned and interrogated
Another key point that the AI tracking industry often fails to mention is GIGO; garbage in, garbage out. The AI platforms are riddled with mistakes, missed mentions, invented answers and various confusions. No data should be taken at face value. It should always be cleaned and interrogated.
- Completeness means all the expected data is present. It prevents blank fields and partial records.
- Uniqueness means each thing is recorded only once. It prevents inflated metrics and double-counting.
- Timeliness means the data is current when you need it. It prevents outdated decisions and stale profiles.
- Validity means each value follows its format rules. It prevents broken inputs and corrupted entries.
- Accuracy means the data matches real-world truth. It prevents operational errors and fake records.
- Consistency means the same fact agrees everywhere. It prevents conflicting data silos.
For example, re: Completeness:
“Let’s take the SUV category as an example. If you ask ChatGPT, “What’s the best SUV?”, you need to repeat that question to understand how often your brand is visible.
“Profound’s method is to run that single prompt once per day for 30 days. That’s not enough.”, writes Brian Stempeck (he’s competing with Profound).
Or re: Accuracy:
“My big worry is with personalization. It seems that AIs personalize to a huge degree with even a small amount of history. As such, I just don’t know whether the same brands ever get recommended the same way to two different people, even across thousands of “anonymized” runs (which is all that a third-party tool could ever do).
“I’m just worried the whole field of AI tracking is built on this oversight, and we don’t know how bad it is…”, writes Rand Fishkin.
What can, and should be, measured
Just because no AI visibility tracker can honestly report all AI responses, it doesn’t mean tracking AI visibility is worthless. When we measure real but narrow activity, and frame the results for what they are, AI visibility tracking can be very valuable.
At New Story Studio we pull-in a large slice of AI platform prompt activity, and highlight the comparison and movement instead. When prompts are fixed, the change over time and the gap to named competitors still carry “directionally useful” signal.
We’ve been here before. Traditional search went through it too, when “search volume” was quoted as gospel for years before the industry learned it was a modelled estimate with wide error bars. AI visibility is at a similar stage now. The AI visibility dashboards look like measurement, so they get treated as measurement, but the qualifications and caveats required for rigorous analysis can get overlooked, treated as inconvenient.
Robust analysis of AI visibility is important, because real decisions and budget now ride on it. AI visibility is increasingly used to justify spend, set strategy, and report progress to boards. Once a figure drives money and decisions, the cost of it being wrong stops being academic. Loose analysis misdirects real resource.
Work with New Story Studio
New Story Studio helps brands make sense of AI search without selling a false sense of certainty. We use these tools where they add genuine value, for tracking movement and spotting where you are absent, and we frame the numbers honestly.
If you want a clear-eyed read on your visibility in AI search, get in touch.

We help a wide range of businesses with their digital marketing challenges.