AI Visibility Scores Are Mostly Noise. Here's the Fix.
New research shows AI visibility rankings shift wildly between runs. Single readings are statistically unreliable. Here's how SMBs should actually measure AI search presence.
A single AI visibility score tells you almost nothing. New research shows these rankings fluctuate significantly run-to-run, meaning one reading is mostly statistical noise. To get a signal you can actually act on, you need repeated sampling across multiple queries and sessions. For most SMBs, that changes how you should think about AI search measurement entirely.
Why does your AI visibility score keep changing?
Because it was never stable to begin with. New research published via Search Engine Journal confirms what anyone who has run these tools back-to-back has suspected: AI visibility rankings shift meaningfully between runs on the same query. The variance is not a bug in your measurement tool. It is baked into how large language models generate responses. If you are making strategy decisions based on a single visibility snapshot, you are making decisions based on noise.
What does the research actually show?
The core finding is that AI-generated search responses have inherent randomness built in. LLMs use a parameter called temperature that controls how deterministic or creative outputs are. Even at lower temperature settings, the same prompt run twice will often produce different citations, different brand mentions, and different orderings. The research found that visibility scores can swing substantially across runs, enough to make a single reading statistically unreliable as a baseline.
This matters because a growing set of vendors is selling "AI visibility scores" as if they were stable rankings you can track week-over-week. Some of these tools run a query once, record whether your brand appeared, and report that as your score. That methodology has a fundamental flaw: you cannot distinguish a real shift in your AI presence from normal model variance without multiple samples.
How many samples do you actually need?
The research suggests you need repeated sampling across many runs of the same query before a visibility number becomes meaningful. The exact threshold depends on how variable the model is on a given topic, but the directional answer is clear: more than one, probably more than five, and ideally tracked over time rather than captured in a single session.
Think about it the way you would think about A/B testing. A sample size of one tells you nothing. You need enough runs to separate signal from variance. For competitive or high-stakes queries, that bar is even higher.
"A single AI visibility reading is like checking your blood pressure once and calling it your baseline. You need the average across multiple readings to know anything real."
For most SMBs, this means resisting the urge to obsess over a weekly score from a single-run tool. The number will move. That movement is mostly not because your AI presence changed.
What actually drives stable AI visibility?
Here is where operators should focus instead of chasing noisy scores. LLMs pull from a few consistent sources when generating responses:
- High-authority third-party mentions. If reputable publications, review sites, and industry directories reference your brand in context, models have more signal to cite you.
- Structured, answer-ready content. Pages that directly answer specific questions in clear language are more likely to be pulled into AI-generated responses. Vague brand messaging is not.
- Consistent entity signals. Your business name, location, category, and key services should appear consistently across your website, Google Business Profile, and third-party sources. Inconsistency creates ambiguity for models trying to resolve who you are.
- Wikipedia and knowledge graph presence. For brands large enough to qualify, these carry significant weight in how models understand and represent you.
None of these inputs are new. They overlap heavily with what has always driven organic search authority. The difference is the output layer: instead of a ranked list, you get a synthesized answer, and whether your brand appears in that answer depends on how confidently a model can resolve your entity.
How should SMBs actually track AI search presence?
Given the noise problem, here is a practical measurement framework:
| Approach | Reliability | Effort | Recommended? | |---|---|---|---| | Single-run visibility tool, weekly | Low | Low | No | | Multi-run sampling (5+ runs per query) | Medium | Medium | Yes, with caveats | | Tracked across multiple tools and queries | Higher | Higher | Yes for competitive categories | | Direct traffic and branded search correlation | High | Low | Always | | Customer-reported "how did you find us" | High | Low | Always |
The cleanest leading indicator for most SMBs is not a visibility score at all. It is whether branded search volume and direct traffic are trending up over time. If AI systems are surfacing your brand more often, people will search for you by name or land directly. That signal is real. A fluctuating visibility score from a single-run tool is not.
For businesses that want to track AI mentions directly, the approach should be:
- Choose 5–10 specific queries that represent how your ideal customer would ask about your category.
- Run each query at least 5 times per measurement period, ideally across more than one AI tool.
- Record whether your brand was mentioned, how it was described, and whether the context was accurate and favorable.
- Track the average mention rate over time, not the single-run score.
This is more work than plugging into a dashboard, but it is the only way to know if your number moved because something real changed.
Does this mean AI visibility tools are useless?
Not entirely. The better tools are aware of this variance problem and are building multi-run sampling into their methodology. Before you pay for any AI visibility product, ask the vendor directly: how many times do you run each query per reporting period? If the answer is once, the score is decorative.
Some tools are genuinely useful for qualitative monitoring: catching inaccurate brand descriptions, spotting competitor mentions in categories you should own, or flagging when AI responses about your space shift in tone. That kind of monitoring has value. Treating the output as a stable rank to optimize against does not.
What we'd actually do
- Stop optimizing to a single-run score. If a tool reports one number without disclosing sampling methodology, deprioritize it. Use branded search volume and direct traffic as your primary AI impact indicators instead.
- Run your own manual sample this week. Pick your 3 most important buyer queries, run each one 5 times in ChatGPT and in Perplexity, and document whether and how your brand appears. That 30-minute exercise will tell you more than most paid tools.
- Focus inputs, not scores. Build more third-party mentions in credible publications, tighten your entity consistency across the web, and create content that directly answers the questions buyers are asking AI systems. These inputs compound. Chasing a noisy score does not.
FAQ
Why do AI visibility rankings change between runs?
LLMs use a randomness parameter called temperature when generating responses. Even on identical queries, the model may surface different sources, citations, and brand mentions each time. This means a single visibility reading reflects that randomness as much as your actual presence, which is why single-run scores are statistically unreliable.
How many times should I run a query to get a reliable AI visibility reading?
The research points to needing multiple runs, at minimum 5 per query, before you can distinguish real signal from model variance. For competitive or high-value queries, more is better. Tracking averages over time across multiple queries gives you a far more actionable picture than any single snapshot.
What is the best way for a small business to monitor its AI search presence without expensive tools?
Run your top 3–5 buyer queries manually in ChatGPT and Perplexity once or twice a month, 5 runs each, and log whether your brand appears and how it is described. Pair that with branded search volume trends in Google Search Console. Together, these give you a reliable signal at zero cost.
Want this running in your business?
The Skool community is where we show the full builds, share the templates, and help you implement. Three tiers, from team training to fractional AI expert.
- Weekly Q&A with Alex and Cameron
- Templates and frameworks you can steal
- Real builds, running in real businesses
More on Marketing AI
Does AI Recommend Your Business? Here's How to Find Out
AI tools like ChatGPT and Perplexity are replacing Google for local searches. Here's how SMB owners can check if AI recommends them, and fix it if not.
AI SEO Tools Small Businesses Actually Keep Paying For
Which AI SEO tools do small business teams actually keep after the free trial? Ranked by real usage, learning curve, and ROI for SMB operators.
What the Anti-AI Beer Ad Tells SMBs About AI Marketing
A beer brand's anti-AI ad went viral by mocking AI-generated content. Here's what SMB owners must know before using AI in their marketing.