Sparkul

Method

How to measure AI visibility (and why most dashboards lie).

A single “AI visibility score” is not a measurement. It’s a number produced by an undisclosed prompt set, run an undisclosed number of times, against engines that answer differently every time you ask. Here’s what a defensible methodology actually looks like.

Grant McClellandFounder, Sparkul AI Agency10 min read

A single “AI visibility score” is not a measurement. It’s a number produced by an undisclosed prompt set, run an undisclosed number of times, against engines that answer differently every time you ask. Here’s what a defensible methodology actually looks like.

Every week another platform offers to tell a business how visible it is “in AI.” The output is usually one number between 0 and 100, a trend line, and a competitor set. It looks like a rank tracker. It is priced like a rank tracker. It is not a rank tracker, because the thing it claims to be tracking does not hold still.

We think most of these dashboards are measuring their own sampling noise and presenting it as performance. That’s a strong claim, so the rest of this is the evidence for it — and what we think honest measurement requires instead.

The score is not a measurement

Start with the object being measured. A Google ranking is a position in an ordered list that, for a given query, location, and device, is broadly reproducible. You can check it twice and get the same answer. That reproducibility is what makes rank a legitimate metric.

AI answers do not have that property, and the size of the gap is not subtle.

In January 2026, SparkToro published research run with Patrick O’Donnell of Gumshoe.ai: 12 brand-recommendation prompts, 600 volunteers, 2,961 runs across ChatGPT, Claude, and Google’s AI Overviews and AI Mode, collected in November and December 2025. Asking the same question repeatedly returned the same list of brands less than 1% of the time. The same list in the same order appeared in fewer than 1 in 1,000 runs.[]

Rand Fishkin’s conclusion was blunt: “Any tool that gives a ‘ranking position in AI’ is full of baloney.”[]

That is the first thing to understand. If a vendor shows you a position — “you rank #4 in ChatGPT for this prompt” — they have taken one draw from a distribution and reported it as a fact. The number will be different tomorrow whether or not you do anything at all.

Why the same prompt gives different answers

It’s tempting to assume the variance is personalization, or A/B testing, or something the vendor could control for. Some of it is. But a meaningful share of it is architectural, and it exists even under conditions where you’d expect perfect repeatability.

In September 2025, Horace He and Thinking Machines Lab published a technical account of why. Running the same prompt 1,000 times at temperature 0 — the setting that is supposed to make output deterministic — against Qwen3-235B produced 80 unique completions. The most common one occurred 78 times. The first divergence appeared at token 103.[]

The cause isn’t randomness in the model. It’s that inference servers batch requests together, batch size varies with concurrent load, and the underlying kernels aren’t batch-invariant — so the arithmetic changes slightly depending on who else happened to be querying the server at the same moment. When the researchers implemented batch-invariant kernels, all 1,000 completions came back identical.[]

The practical translation: part of your “AI visibility” fluctuation is a function of how busy someone else’s GPU was. That is not a signal about your brand. No amount of dashboard polish converts it into one.

Every engine is a separate instrument

The second lie in the single-score model is aggregation. Blending ChatGPT, Gemini, Perplexity, and Google AI Mode into one index assumes they’re measuring the same thing. They aren’t.

They differ at the retrieval layer. Perplexity is retrieval-first by design — query, live search, synthesize, cite. Google’s generative surfaces are wired into its live index and Knowledge Graph. ChatGPT historically answered substantially from model-internal knowledge unless it decided to search.[] Those are different systems with different inputs. A signal that moves one may not touch another.

The measured behavior bears this out. An April 2026 arXiv study by Julius Schulte, Malte Bleeker and Philipp Kaufmann tracked four engines across four verticals for 45–46 days. Between two consecutive days, only about 35% of cited sources overlapped — meaning roughly 65% of the sources an engine used changed overnight. Brand overlap was steadier but still moved, with Jaccard similarity between 0.45 and 0.59.[]

There is also the reader-facing question of whether a citation is even visible. Search Engine Land’s July 2026 analysis of 16 million brand appearances found that roughly 40% of AI citations did not include the brand’s name in the answer — with rates ranging from 52% on Perplexity and 49% on Google AI Mode down to 19% on Microsoft Copilot. Your content gets used; your name doesn’t appear.[] As the authors put it, “a citation your reader never sees is visibility that only exists in a dashboard.”[]

If citation rate, mention rate, and reader-visible attribution all differ by engine, then averaging them produces a number that describes nothing. Engines get tracked separately or they don’t get tracked.

How many runs is enough

This is the question almost no vendor answers in public, and it’s the one that determines whether a number means anything.

The Schulte et al. paper ran up to 10 repeated queries per prompt on the same day and found that intra-day variation alone accounted for most of the observed instability: same-day source similarity averaged a Jaccard of 0.32–0.43, brand similarity 0.33–0.48.[] In other words, most of the movement you’d be tempted to read as a trend happens within a single afternoon.

Their sampling analysis is the most useful number in the entire literature. A single run produces a standard error of 0.370 — wide enough that a true per-brand detection rate of 50% could be observed anywhere from -22% to +122%. To get standard error below 0.10, they found you need at least 7 runs per prompt per day for brand detection, and at least 8 for source coverage. Even then the 95% confidence interval is roughly ±0.158. To stabilize a per-brand detection rate you need rolling aggregation over two to four weeks: at 10 days the standard error is about 0.107, at 21 days about 0.053.[]

Read that again, because it settles the argument. One run of one prompt is not a weak measurement. It is not a measurement.

SparkToro’s researchers ran each prompt 60–100 times per platform and still flagged the statistically sound sample size as an open question.[] That is what intellectual honesty looks like in this space, and it is the opposite of a confidence-inspiring score out of 100.

The query universe is where most of the lying happens

Sampling is a solvable engineering problem. Prompt selection is a harder, more interesting one, and it’s where vendor theater is easiest to hide.

Every AI visibility product runs some list of prompts. That list is the entire experiment. Change it and the score changes. Almost none of them publish it.

The problem is worse than opacity, because real people don’t ask questions in a standard way. SparkToro analyzed 142 human-written prompts on the same topic and found an average semantic similarity of 0.081 — near-total variation in how humans phrase the same underlying need.[] A curated set of fifty tidy prompts written by a marketer is not a sample of that space. It’s a sample of how marketers write prompts.

And unlike search, there is no keyword volume data for prompts. Nobody publishes how many people asked ChatGPT a given question last month. Any vendor that shows you prompt-level “volume” is presenting a model as an observation, and you should ask, specifically, what it was modeled from.

Then there’s personalization. OpenAI has offered memory since February 2024, explicitly designed so that ChatGPT carries details between conversations and applies them to later ones — “Remembering things you discuss across all chats saves you from having to repeat information and makes future conversations more helpful.”[] A logged-in user with history is not being served the same system as a clean API call. Both are legitimate things to measure. They are not the same measurement, and a defensible methodology says which one it ran.

A query universe we’d defend has four properties: it is written from actual customer language rather than invented by the agency; it is explicitly segmented by intent, because a research question and a purchase question behave differently; it is versioned, so that a score change can be checked against a prompt-set change; and it is visible to the client in full. If you can’t see the prompts, you can’t audit the number.

What analytics can and cannot see

The measurement problem has a second half, and it is harder than the first: figuring out what any of this produced.

Referral traffic is the obvious candidate and the weakest one. Semrush’s analysis of 50,000+ websites across 17 industries for calendar 2025 found AI traffic growing 66% year over year — from 462 million to 767 million monthly visits — while still accounting for 0.14% of total visits, against 16.04% for organic search. Google AI Mode alone went from roughly 1,600 visits in January 2025 to 38.2 million in December, which is 0.01% of the total.[]

Both facts are true at once: the fastest-growing channel, and a rounding error in volume. Using referral sessions as your primary AI visibility metric means grading yourself on the thinnest available slice of the phenomenon.

It’s also a slice with holes in it. In May 2025, Search Engine Land documented that links in Google’s AI Mode carried a noreferrer attribute stripping the referrer value, so visits landed in analytics as Direct or Unknown and clicks did not surface in Search Console. Lily Ray called it “Not Provided 2.0.” Google’s John Mueller said it looked “unexpected,” and the article was updated a week later to note Google had fixed it.[] The episode matters less as a scandal than as a demonstration: your attribution depends on a platform’s referrer policy, which can change without notice, in either direction.

Google’s own reporting has a defined ceiling. When generative AI performance reports arrived in Search Console in June 2026, they included impressions, pages, countries and devices — and no click data at all. AI Mode and AI Overviews are aggregated together with no breakdown by feature, and the rollout began with a subset of site owners in the UK.[] Asked about the missing clicks, a Google spokesperson said only that additional metrics would come “over time.”[]

So here is the honest ledger. You can measure, with real rigor: how often your brand appears in answers to a defined prompt set, per engine, over repeated runs; which sources those engines cite; whether your name appears alongside the citation; and the referral sessions that do carry an identifiable referrer, with their downstream conversion behavior.

You cannot measure: total prompt volume for your category; the number of people who saw your brand in an answer and never clicked; the revenue influenced by an answer that produced no session; or a clean causal link between a content change and a visibility change on a system that changes 65% of its cited sources overnight anyway.[]

Anyone selling you certainty on the second list is selling you a story.

Telling methodology from theater

Five questions. Ask them of any vendor, including us.

How many times do you run each prompt, and over what window? If the answer is fewer than seven runs per prompt per day, or if it’s a single point-in-time crawl, the confidence interval swamps the finding.[] Digiday’s May 2026 reporting found that many tools deliver exactly that — “point in time” results rather than ongoing measurement.[]

Can I see the full prompt set, and its version history? No list, no audit.

Are engines reported separately, with per-engine methodology? A blended score is a category error.

Do you distinguish citation from mention? Roughly 40% of citations don’t carry the brand name; a tool that counts them identically is inflating your result by construction.[]

What do you publish about your own error rate? This is the tell. Serious measurement organizations publish their limits. The Tow Center at Columbia tested eight generative search tools with 1,600 queries and reported that the tools returned incorrect answers to more than 60% of them, with per-engine rates from 37% for Perplexity to 94% for Grok 3, and noted that ChatGPT signaled uncertainty only 15 times across 200 responses while getting 134 wrong.[] A vendor that never publishes a confidence interval is telling you something about their confidence intervals.

There is a market reason this matters. Digiday reported agencies paying up to $1,000 a month for these platforms, with Paul Dyer, CEO of /prompt, observing: “If you use three different tools and give them the same prompts, you get three different answers.”[] Ryan Mason of Markacy described the tools as benchmarking instruments rather than sources of truth.[] That’s a fair use of them. It is not what the dashboards imply.

The stance

AI visibility is measurable. It is measurable as a distribution, not a position — a share of appearances across many prompts and many runs, reported per engine, with a stated confidence interval and a stated prompt set. Fishkin’s formulation is the right one: “Visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric.”[]

Everything narrower than that — a score, a rank, a single crawl, a blended index — is a compression that destroys the only information worth having.

We would rather give a client a wide interval that is true than a tight number that is invented. That means some months the honest answer is “this moved, but not beyond the noise floor,” and some questions get the answer “we cannot know that, and neither can anyone selling you a tool that says otherwise.”

That is not a weaker offer. It is the only version of this work that survives contact with the data.

Sources

  1. [1]NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. Rand Fishkin & Patrick O’Donnell. SparkToro (with Gumshoe.ai). January 28, 2026. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
  2. [2]Defeating Nondeterminism in LLM Inference. Horace He. Thinking Machines Lab. September 10, 2025. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
  3. [3]Don’t Measure Once: Measuring Visibility in AI Search (GEO). Julius Schulte, Malte Bleeker, Philipp Kaufmann. arXiv preprint 2604.07585. April 8, 2026. https://arxiv.org/abs/2604.07585
  4. [4]AI Search Has a Citation Problem. Klaudia Jaźwińska & Aisvarya Chandrasekar. Columbia Journalism Review, Tow Center for Digital Journalism. March 6, 2025. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
  5. [5]Marketers question expensive AI visibility tools as inconsistent results fuel skepticism. Kimeko McCoy. Digiday. May 1, 2026. https://digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/
  6. [6]Google’s AI Mode traffic is untrackable. Danny Goodwin. Search Engine Land. May 22, 2025 (updated May 28, 2025). https://searchengineland.com/googles-ai-mode-traffic-untrackable-455883
  7. [7]Google Search Console AI performance reports and controls to block your content in AI responses. Barry Schwartz. Search Engine Land. June 3, 2026. https://searchengineland.com/google-search-console-ai-performance-reports-and-controls-to-block-your-content-in-ai-responses-479298
  8. [8]We analyzed billions of web visits: How AI is reshaping traffic channels. Margarita Loktionova. Semrush. April 27, 2026. https://www.semrush.com/blog/traffic-channel-mix-study/
  9. [9]Ghost citations: Why AI search cites your content, not your brand. Nikki Lam & Samanyou Garg. Search Engine Land. July 29, 2026. https://searchengineland.com/ghost-citation-problem-ai-483794
  10. [10]How different AI engines generate and cite answers. Greg Jarboe. Search Engine Land. October 10, 2025. https://searchengineland.com/how-different-ai-engines-generate-and-cite-answers-463234
  11. [11]Memory and new controls for ChatGPT. OpenAI. February 13, 2024. https://openai.com/index/memory-and-new-controls-for-chatgpt/

Want to see how your business shows up in AI search?

Run our free AI Visibility Scorecard. We’ll show you which AI engines are mentioning you, which competitors are winning the queries that matter, and exactly what’s missing from your visibility setup.

Run your free Scorecard →

More from Sparkul