Last updated August 2026
AI Search visibility tools give marketers wonderfully precise-looking numbers: 67% visibility, 90% visibility, position #1, 100% share of voice.
How much confidence should you put in those numbers?
I had an unusual opportunity to find out. I currently have access to several AI Search visibility platforms, so I tracked the exact same prompt for the exact same brand across six tools. All the results below were collected within less than one week.
The tools often agreed that Marketing With Dave was visible. They did not agree on how visible it was, where it ranked, or even whether it appeared at all on a given AI platform. One tool gave me a 90% visibility score while the underlying AI response got my name wrong.
How I Ran the Test
I used one of the 10 prompts I benchmark AI Search visibility platforms with:
Who is Marketing With Dave?
This is intentionally a branded prompt. It isn’t a difficult query for Marketing With Dave to appear for, which is exactly what makes it useful for comparing how different tools measure the same basic outcome.
I pulled results from Nuwtonic, SnowSEO, Subsig, Visby, ZeroRank, and iGEO. iGEO is included because it was part of the test, although it never returned any results during the testing period.
This isn’t a scientific study. The tools may query different model versions, run prompts at different times, and use different sampling methods, locations, or scoring formulas. AI responses themselves are also non-deterministic. That’s not a weakness in the experiment. Those differences are part of the measurement problem. I also didn’t repeatedly rerun every tool to determine how much these results would change from one run to the next. That’s an important limitation, but also part of the larger problem: a single visibility measurement may not be particularly repeatable.
That uncertainty isn’t theoretical. SparkToro and Gumshoe recently ran nearly 3,000 AI responses and found substantial inconsistency in brand recommendations even when prompts were repeated. Their research tested the stability of the AI responses themselves. My test looks at another layer of the problem: whether the commercial tools measuring those responses agree on what the measurement means.
The Same Prompt, Four Different Scores
Here’s what the tools reported for the same prompt on OpenAI alone:
| Tracking Tool | Mentioned | Position | Visibility |
|---|---|---|---|
| Nuwtonic | Yes | #1 | 100% |
| SnowSEO | Yes | #1 | 75% |
| Subsig | Yes | #1 | 16.7% |
| Visby | Yes | — | 90% |
| ZeroRank | Yes | #2 | Not shown separately |
One platform, one prompt, one week. Four visibility numbers: 100%, 90%, 75%, 16.7%. That doesn’t mean three of the four tools are wrong. It means a number labeled “Visibility” isn’t automatically measuring the same thing from one product to another.

Semrush, one of the largest companies in the SEO space, defines AI Visibility in its Visibility Overview as a 0-100 benchmark comparing how often a brand appears against competitors, but defines visibility differently inside Prompt Tracking, where it reflects a site’s standing within the top citations for specific tracked prompts. Related concepts, not identical measurements. Semrush also warns that AI responses are fast-changing and personalized, so no platform can produce exact visibility numbers, only directional signals. Semrush explains where its AI visibility data comes from here.
That’s also why I question the precision implied by some of these dashboards. A visibility score reported as 16.7% looks remarkably exact for a measurement built on AI responses that may change the next time the prompt runs.
That context matters when a dashboard hands you a 96% or a 67% without making its methodology equally visible.
100% Visible in Gemini. Or 0%. Or, Actually, It Depends What You Mean By Gemini.
The most dramatic disagreement in the data isn’t about scoring formulas. It’s about whether Marketing With Dave appeared at all, and it gets stranger the closer you look.
Nuwtonic reported the Gemini result as mentioned, position #1, visibility 100. Visby reported the same prompt on Gemini less than a day later as not mentioned, visibility 0%.
That’s not a rounding error.
But Visby’s own dashboard adds a wrinkle worth sitting with: alongside that 0% Gemini score, Visby separately tracks Google’s AI Overview, and there it reports Marketing With Dave at 100% visibility with 7 citations. Same brand, same week, same company (Google), two different surfaces, opposite results, from a single tool that at least has the honesty to report them separately.
Gemini the chatbot and AI Overview the search feature are genuinely different products built on related but distinct retrieval behavior. A visibility score that doesn’t specify which one it’s measuring, or that quietly folds both into “Google,” isn’t just imprecise. It’s answering a different question than the one you think you’re asking.
A 90% Visibility Score Can Still Be Wrong
Visby gave Marketing With Dave a 90% OpenAI visibility score. The answer cited MarketingWithDave.com and correctly connected the site to digital marketing, A/B testing, martech, measurement, and SiteSqueeze.
It also said Marketing With Dave was run by Dave Harrell.
It isn’t.
So I could check off every box that usually signals success: brand mentioned, high visibility score, website cited, relevant topics identified. And the answer still contained a material entity error. What good is a visibility score if the AI doesn’t accurately understand the entity it’s making visible?
Marketing With Dave probably makes this problem easier to spot than most brands would. Dave is a common name, “marketing” is a common word, and there’s no shortage of marketers and agencies using similar language. A highly distinctive SaaS brand may resolve more cleanly. But that’s exactly why the ambiguity doesn’t invalidate the test: real companies share names, terminology, and categories constantly. Entity resolution is part of AI visibility, and a high score can hide a wrong answer.
Even One Tool Isn’t Internally Consistent
You don’t need to compare across vendors to find volatility. Nuwtonic’s own results for this single prompt run remarkably tight across five platforms: 95 to 100 visibility and position #1 on Perplexity, OpenAI, Gemini, Grok, and Meta. Then Anthropic drops to 80 and position #5.
Same tool, same prompt, same week, same methodology. The one variable that changed is the model being queried. If a single vendor’s own numbers can swing that much platform to platform, treating any one visibility score as a stable fact about your brand, rather than a snapshot of one model on one day, is asking more of the number than it can deliver.
The Tools Don’t Even Report the Same Things
The differences go beyond how visibility gets calculated. The six products vary considerably in what they expose to the user at all:
| Tool | Visibility | Position | Share of Voice | Sentiment | Citation Count | Cited Pages |
|---|---|---|---|---|---|---|
| Nuwtonic | ✓ | ✓ | – | – | ✓ | ✓ |
| SnowSEO | ✓ | ✓ | ✓ | ✓ | ✓ | ⚠ |
| Subsig | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Visby | ✓ | – | – | – | ✓ | ✓ |
| ZeroRank | ✓ | ✓ | – | ✓ | ✓ | ✓ |
| iGEO | – | – | – | – | – | – |
More metrics isn’t automatically better; it can just mean more numbers to misread. But the gaps matter once you start comparing products or trying to build a repeatable process.
Mentions, Citations, Accuracy: Three Different Questions
A brand can appear in an AI answer without its website being used as a source, and a page can get cited without the brand being prominently recommended. Semrush draws the same line: mentions are appearances of the brand in AI answers, citations are linked references to the content behind them. Semrush’s AI Share of Voice guide explains that distinction.
I’d add a third category: accuracy. Did the AI get the brand, the person, and the positioning right? That may matter more than whether your visibility score moved from 67 to 72, and Visby’s citation counts for this one prompt ranged from 0 (Gemini) to 7 (AI Overview) without any change in whether the brand was “visible” by the tool’s own definition. Citation volume and visibility score aren’t the same thing either.
Platform-Level Data Beats a Blended Score, and I Have Proof
ZeroRank reports an “Overall, 3 models” figure alongside its per-platform numbers: 67% visibility, position #1.5. It’s a real, sensible-looking average. It also sits neatly between the individual platform numbers and, in doing so, erases the exact volatility this entire test was built to surface: a #1 on Perplexity, a #2 on OpenAI, and no mention at all on Google/Gemini, blended into one tidy 67.
Semrush reaches a similar conclusion in its own methodology, reporting AI platforms separately for cross-platform comparisons rather than combining them into a single weighted score, because aggregation can hide insights. See Semrush’s AI Visibility Index methodology. After watching one tool’s blended score paper over a #1-to-absent spread within its own data, that conclusion doesn’t need much convincing.
What I Actually Trust
Not any single visibility score. Before I’d act on one, especially a low one, I want to know:
- Do I already have a page that genuinely answers this prompt, or is this a content gap rather than a visibility problem?
- Does the result hold up if I check again next week?
- Does the brand show up on more than one platform, or is this one model’s quirk?
- Do multiple tools tell a broadly similar story?
- Which specific pages are actually getting cited?
- Are competitors consistently showing up where I’m absent?
- Is the AI’s description of my brand actually correct?
I’d apply a similar standard when choosing the software itself. If a platform won’t let me see the underlying AI response, platform-level results, and citations behind its score, I’m much less interested in the score. I want enough raw evidence to question the metric, not just a dashboard asking me to trust it.
That first question turned out to be the most useful one in this whole exercise. Some of the prompts I want Marketing With Dave to appear for don’t yet have an obvious page on the site that deserves to be cited. That’s not an AI visibility problem. That’s a content gap, and it calls for a different response than chasing a better score.
Bottom Line
AI Search visibility measurement is useful, and I’m going to keep using it. But the problem isn’t that these tools are wrong. It’s that their precise-looking scores can imply a level of standardization and certainty the underlying measurements don’t support yet, whether that’s one vendor’s 0% sitting next to another vendor’s 100% for the same platform, or one vendor’s own results shifting from position #1 to #5 depending on the model.
The question isn’t simply “What’s my AI visibility score?” It’s “What decisions does this measurement actually justify?”
Treat AI visibility scores as directional evidence, not ground truth. The real work isn’t chasing a perfect number. It’s using repeated measurements, actual AI responses, citations, competitor patterns, and content gaps together to figure out where your brand is genuinely becoming part of the answer, and right now the industry is much better at producing these scores than at explaining how much confidence you should put in them.