Answer engines are non-deterministic. Understanding what that does to your evidence is the core skill this course exists to teach.
The number moves on its own
Ask the same prompt twice with nothing changed and you can get different sources, a different set of competitors named, and sometimes a different recommendation. The output is a draw from a distribution.
This means your measurement has natural variance independent of anything you do. If you observe a change after shipping something, part of that change is your work and part is the distribution. With a single before-and-after observation you cannot tell which part is which — and the honest answer is often that you cannot tell at all.
How this goes wrong in practice
The failure is not usually dishonesty, it is asymmetric attention. A team ships a change, sees the number rise, and stops looking. The same team ships a change, sees the number fall, and concludes the measurement is noisy. Both readings might be right, but applying them selectively guarantees the wrong conclusion.
The defence is deciding in advance. Before you re-measure, write down what movement you would count as success and what you would count as failure. Then hold yourself to it even when you dislike the answer.
Some prompts are noisier than others
Variance is not uniform. Some questions have a settled answer that the engines return consistently; others are contested, with the named set shifting between draws.
This distinction is directly useful. A settled prompt where you lose is expensive to move but the win is durable. A contested prompt is cheaper to move and less stable — and it is also where an opening still exists, because the consensus has not formed yet. Knowing which of your prompts are which tells you where effort compounds.
Check yourself
If your number moves after a change, what are the two possible explanations?
What movement would you count as failure? Write it down before you look.
Which of your prompts do you expect to be settled, and which contested?
Go deeper
This lesson is the spine. These guides are the depth.
A single AI check is one sample, not a measurement. How many repeated runs and prompts you need before an AI-visibility number is trustworthy, with the published evidence on sample size and confidence intervals.
AI visibility scores move week to week even when nothing you control has changed. Here are the 6 sources of normal variance, the patterns that signal real movement, and what to watch.
Volatile AI-citation queries aren't a maintenance burden. They're contested open slots no source has won yet. How to read variance as a query-selection signal.
The contested-vs-settled signal tells you which AI queries are winnable. This is how to act on it: measure the noise floor per query, classify each query's phase, route each phase to a different move, and avoid the trap that wrecks the whole thing, mistaking a query that's dying for one you're losing.
AI answers are non-deterministic, so a naive GEO test lies to you confidently. The four test-design rules that decide whether your AI-visibility result is signal or noise: measure a rate, test at the buyer's specificity level, test per language, and sort each cited source by the move.
Turn AI answer gaps into governed marketing execution.
Start free with a ChatGPT visibility read, then add multi-engine tracking, Corporate Context governance, and the action-to-result loop when you are ready.