SolCrys Logo

Citation & Source Influence

Contested vs Settled: Your Most Volatile AI Queries Are the Open Slots

AI-citation volatility is a two-regime signal: contested queries reshuffle because no source has won the canonical slot yet, while settled ones are stable because one has. The high-variance, high-value queries most teams write off are the open slots still up for grabs. Published research shows why you can't judge them from one run: repeated runs of the same prompt within 24 hours overlap by only about a third in the sources they cite, so you earn your way in by measuring presence as a rate over repeated runs, per engine, rather than reacting to a single reading.

By Eason Wang, Co-Founder & CPO, SolCrys

Published · Updated

Questions this guide answers

  • Why do my AI citations change every week?
  • Which AI queries should I target first for GEO?
  • Are volatile AI-citation queries worth optimizing?
  • What is a contested vs settled AI query?
  • How do you win an unstable AI-citation query?

Direct answer

AI-citation volatility is a signal about the query, not a flaw in your tracking. Some queries reshuffle their cited sources run after run; others sit rock-stable for weeks. The volatile ones are not a maintenance burden to be quieted down. In our reading they are the open slots: queries where no source has yet won the canonical position an engine reaches for, so the answer keeps drawing from a shifting handful of candidates.

We split the spectrum into two regimes. A contested query is unbranded discovery work, the "best X for Y" and feature-comparison questions where the engine has no settled go-to source and the cited set churns. A settled query is branded, factual, or definitional, where one canonical source has won and the answer barely moves. The reframe that follows is the whole point of the page: a high-value query that looks volatile is usually a slot still up for grabs, not a write-off.

You earn your way into a contested slot by treating it as a process, not a single reading. Measure presence as a rate over repeated runs instead of reacting to one run. Separate ordinary sampling run-variance from genuine week-over-week drift. Then feed consistent, fresh, human-approved signals across the independent sources each engine samples, and re-measure to see whether the rate is climbing. We keep three kinds of claim separate throughout: the regime split and the settling mechanism are our labeled interpretation; the run-variance discipline rests on published research; and the worked example is an illustrative scenario, not client data.

Why citations move at all

Before treating variance as a selection signal, it helps to know why an AI-citation number moves even when nothing you control has changed. The mechanics live in why your AI visibility score moves: engines are non-deterministic by design, models get updated, competitors publish, third-party citations shift, freshness windows turn over, and sampling cadence introduces its own jitter. A score is a statistical estimate, not a deterministic count, so some movement is normal noise.

Published research puts a size on that noise. SparkToro's January 2026 study ran 12 prompts 2,961 times across ChatGPT, Claude, and Google's AI Overviews and found less than a 1 in 100 chance that two runs return the same list of brands, and less than 1 in 1,000 that they return it in the same order. Cited sources move as well: in our own repeated runs, the set of sources cited for the same prompt changes from day to day.

The score-moves page draws one conclusion from this: the moving score is a symptom, and the diagnostic job is to ask what was published and which citations changed, rather than panic at the number. This page makes the opposite move on the same raw phenomenon. Here, variance is a selection signal: a map of which queries are still contested and therefore still winnable. Same noise, inverted use.

The two regimes, and how to tell them apart

Contested and settled are the two ends of a spectrum, not a hard binary, and most prompts sit somewhere between. The diagnostic is not a single threshold; it is a pattern you read over several runs. A settled query shows high, stable presence with a small, recurring set of cited sources. A contested query shows low or middling presence with a cited set that reshuffles run to run, often pulling in a different domain each time.

Illustrative scenario only. Picture a mid-market observability vendor, call it Brand O, that tracks 30 prompts on 4 engines and collects 32 responses per prompt each week (8 runs on each engine). Its settled end is a branded lookup, "What does Brand O do?": Brand O appears in 31 of 32 responses, and the same 2 sources, its own documentation and one analyst profile, are cited in most of them. The engines already have a place to go. Its contested end is a discovery prompt, "best log management tool for Kubernetes teams": Brand O appears in 4 of 32 responses, Datadog, Grafana Labs, and Elastic take most of the named slots, and the cited set pulls in a different comparison post or community thread on nearly every run. No source has locked that slot, including the leaders' own pages.

One branded prompt and one discovery prompt illustrate the two ends. They are not a measured finding that branded or factual queries are stable as a class. The practical tells for sorting a query yourself:

  • Presence level: settled queries cluster near the top (consistently high presence); contested queries sit low or in the middle and bounce.
  • Source stability: settled queries cite the same 1 or 2 sources every run; contested queries pull a different domain into the answer from run to run.
  • Query type as a prior, not a rule: branded, factual, and definitional questions skew settled; unbranded "best X for Y" and feature questions skew contested. Treat type as a hint to check, never as the verdict.
  • Read it over runs, not once: a single run cannot tell the two apart, because one run of a contested query can look deceptively stable by luck. The regime shows up in the pattern across several runs.

The reframe: a volatile high-value query is an open slot, not a write-off

Here is the inversion, stated plainly as our interpretation. Most teams see a high-value query that swings between runs and conclude it is unstable, unreliable, or not worth the effort, and they redirect attention to queries that already read steady. That instinct quietly selects for slots someone else has already won, where your upside is smallest.

We read the same volatility the other way. A contested slot is churning precisely because the position is unclaimed. The engines are still shopping among candidate sources because none has earned the canonical spot. That is the moment a new source can become one of the corroborating few. A settled slot, by contrast, is stable because the contest already ended, and entering it means dislodging an incumbent the engines have learned to trust. So variance flips from a symptom you suppress into a selection signal you act on: the volatile, high-value queries are the ones to prioritize, not the ones to write off.

This is an interpretation, not a measurement, and we want to be exact about the boundary. No study we know of proves that prioritizing volatile queries produces citations; a July 2026 survey of 45 GEO studies found that "no reviewed technique shows a stable, longitudinal, cross-platform causal effect" on discoverability (Martinez, 2026). The claim is that the logic is sound and worth testing on your own prompt set: target the high-value queries that are still contested, because those are the slots still open. The broader "design a test you cannot fool yourself with" frame lives in the testable GEO playbook, and turning the variance into a spend decision is covered in contested, settling, or decaying.

Run-variance vs non-stationarity: the discipline to hold yourself to

The reframe only works if you can tell a query that is genuinely moving from one that is merely sampling noisily. This is the part the evidence supports most firmly:

  • Cited sources churn. In our own repeated runs, the set of sources cited for the same prompt changes from day to day, so one day's source list is not a baseline.
  • Brand lists churn too. Two runs had less than a 1 in 100 chance of returning the same list of brands (SparkToro, 2026). The same study calls visibility measured as a rate across prompts run multiple times "a reasonable metric" and dismisses "ranking position in AI."
  • Sample size decides what you can see. One run is one draw, and on a small sample a single mention appearing or vanishing moves the rate by several points, so a presence number needs many runs per prompt, per engine, before a change in it means anything.

What this means in practice

Illustrative scenario only. Go back to Brand O's contested prompt. At 32 responses a week, one mention appearing or vanishing moves presence by about 3 percentage points on its own. Over 6 weeks with no content change, Brand O's weekly count on that prompt reads 4, 3, 5, 4, 2, and 5 of 32, roughly 6% to 16%. That looks dramatic on a chart, and it is nothing: the series oscillates inside a band and goes nowhere. A week-over-week jump from 2 to 5 is the kind of move teams publish as a win, and it sits entirely inside the noise.

The rule that falls out of this is simple. Measure presence as a rate over weeks of runs, never from a single reading and never from a 2-run difference. Run-variance is the wobble you see when nothing changed; non-stationarity is a sustained move that survives many runs. Only the second is a real signal. For the statistics on how many runs it takes before a presence number is trustworthy, see how many runs until your AI-visibility number is trustworthy.

How a contested slot settles: per-engine corroboration

If a contested slot can settle, what does the settling look like? Our working hypothesis, labeled as one on purpose, is that a slot settles per engine when several independent sources corroborate the same claim, giving the engine a stable place to reach. The full mechanism is in AI cites consensus, not authority, which argues that engines repeat the claim corroborated across independent sources rather than ranking a single authoritative page. "Agreement settles the slot" is the one-line version, and it is a hypothesis, not a measured result.

Why per engine? Because the engines read different slices of the web. A SIGIR 2026 study found that Google Search, AI Overviews, and Gemini retrieve substantially different sources, with average Jaccard similarity below 0.2 (Grossman et al.). If each engine samples a different handful of sources, a slot has to settle separately inside each one as its sampled sources start to agree. That is consistent with the hypothesis without proving it; nothing in the overlap data demonstrates the causal step from agreement to a settled slot.

Settled is also not permanent, and the clearest 2026 example had nothing to do with content. In the second week of August 2026, ChatGPT all but stopped citing Reddit: independent trackers reported by Search Engine Journal saw Reddit's share of ChatGPT citations fall from 3.83% to 0.52% in about a week. In our own tracking, Google AI Overviews, Perplexity, and Gemini barely moved. Slots that looked settled on ChatGPT reopened on that engine alone, because of a change on ChatGPT's side rather than anything a brand published.

Three consequences follow. First, there is no blanket "AI engines do X" here; a slot can be settled in Claude and still contested in Perplexity, so you read and act per engine. Second, settling is not instant. New material has to be crawled, indexed, and corroborated, which takes days, not minutes; the clock is daily, not real-time. Third, re-check settled slots too, because an engine's retrieval policy can reopen them without warning.

How to earn a contested slot, governed

Earning into a contested slot is the same governed loop SolCrys runs on everything else: Measure → Diagnose → Execute → Verify. Nothing about it is automatic, and the Execute step stays human-approved. We do not claim the engine publishes for you, and we do not promise a slot; you earn your way in, and the engines decide.

The loop applied to a contested query looks like this:

  • Measure the rate. Track presence for the query as a rate over weeks of runs, per engine, not from one reading. A single run cannot tell you whether a slot is contested or settled, and a 2-run difference is usually just run-variance.
  • Diagnose contested vs settled. Use the tells above: presence level, source stability across runs, and query type as a prior. Confirm the high-value queries you care about are genuinely contested (low and churning), which is what makes them open rather than already lost.
  • Execute consistent fresh signals across sampled sources, human-approved. Because each engine samples a different handful of independent sources, the work is getting your claim corroborated across several of them over time, not perfecting one authoritative page. Every piece is reviewed and approved by a person before it ships, and none of it is a planted mention; Google's own guidance says seeking inauthentic mentions "isn't as helpful as it might seem."
  • Verify on a re-run. Re-run the same frozen prompt set on the same engines and look for a sustained climb that survives many readings, the non-stationarity signal rather than a one-run blip. A before-and-after is an observation, not proof of cause. If the rate holds its rise across weeks, the slot is settling for you; if it wobbles inside its band, nothing real has moved yet.

Where owned content fits, and where it does not

Your own domain is one corroborating source among several, and its weight now differs by engine. On ChatGPT, since August 2026, official vendor sites, documentation, and government pages have taken much of the share Reddit lost in our tracking, so a clear, current canonical page counts for more there than it did. That is a gain for official pages as a class, not a guaranteed gain for yours; your docs now compete directly with your competitors' docs. On Google's AI surfaces, Google says optimizing for its AI features is "still SEO," so the page has to earn its place in Search first. Either way, owned content earns the click when an engine surfaces you, and on contested discovery queries, third-party corroboration is usually what earns the mention.

One more sequencing note: do not delete the contested prompts that read 0% today. A query with no presence yet is an open slot you have not entered, not a dead one. The four-state way to manage a prompt set so you keep these candidates instead of pruning them is laid out in the AEO prompt-set lifecycle.

What we don't claim

To keep the method honest, here is the boundary of what this page does and does not assert. The variance discipline rests on published research. The interpretation above it is ours, offered as a frame to test on your own prompt set, not as a proven law.

  • We do not claim branded or factual queries are stable as a class. The settled example is one illustrative branded prompt. It shows what a settled slot looks like; it is not a measured cohort.
  • We do not publish small run-to-run moves as slots being won or lost. When one mention moves presence by several points, a change inside that band is sampling run-variance, not a trend.
  • We do not claim a slot settles because sources agree. That is a hypothesis the cross-engine overlap research is consistent with, not a measurement. The causal step from corroboration to a settled slot is unproven, and we label it a hypothesis everywhere it appears.
  • We did not run a study ranking many query types by citation stability. The regime framing is a logical reading of how citations behave, generalized as method. No such study sits behind this page.
  • We do not promise a slot, or claim the engine publishes for you. Contested slots are up for grabs; you earn your way in with human-approved content, and the engines decide. The settling clock is daily, not real-time.
  • We measure per engine, not as one blob. Engines retrieve substantially different sources, and a retrieval change on one engine can reopen slots there alone, as ChatGPT's August 2026 shift did. There is no blanket "AI engines do X" claim here.

Start with your own contested slots

The fastest way to see which of your high-value queries are contested and which are settled is to measure them as a rate across runs and read the variance as a selection signal. The free SolCrys workspace starts on ChatGPT: 10 prompts and 1 content audit, free, no credit card, with email verification. Treat the first reading as a baseline, not a verdict; tracking the other engines is on paid plans. Start Free.

Sources

FAQ

Why do my AI citations change every week?

Part of it is normal noise and part of it is the query type. Engines are non-deterministic, models update, competitors publish, and sampling cadence adds jitter, so a citation number moves even when nothing you control changed. The noise is large: SparkToro's January 2026 study of 2,961 runs found less than a 1 in 100 chance that two runs return the same list of brands, and in our own repeated runs the set of sources cited for the same prompt changes from day to day. On top of that, some queries are contested: no source has won the canonical slot yet, so the cited set genuinely reshuffles. Engines also change retrieval policy; ChatGPT all but stopped citing Reddit in August 2026. Measure presence as a rate over weeks, per engine, before reading any week's move as real.

Which AI queries should I target first for GEO?

Target the high-value queries that are contested rather than already settled. A contested query shows low or middling presence and a cited set that reshuffles run to run, which means no source has locked the slot and the position is still open. A settled query shows high, stable presence with the same 1 or 2 sources cited every run, which means the contest already ended and entering it requires dislodging a trusted incumbent. Read each query's regime over several runs, not from one reading, and prioritize the volatile high-value ones because those are the open slots. This prioritization is our interpretation, offered as something to test on your own prompt set.

Are volatile AI-citation queries worth optimizing?

Often yes, and that is the reframe at the center of this method. Most teams write off a volatile high-value query as unstable and chase queries that already read steady, which quietly selects for slots someone else has already won. We read the same volatility as a selection signal: a contested slot churns precisely because the canonical position is unclaimed, which is the moment a new source can become one of the corroborating few. The caveat is to first confirm the movement is real and not sampling noise. With a few dozen responses per prompt, one mention flipping moves presence by several points on its own.

What is a contested vs settled AI query?

They are the two ends of a spectrum, not a hard binary. A settled query is one where a canonical source has won the slot, so the answer is stable, presence is high, and the same sources are cited every run; branded, factual, and definitional questions skew this way. A contested query is one where no source has won, so presence is low or middling and the cited set reshuffles from run to run; unbranded discovery and feature questions skew this way. Treat query type as a prior to check, not a verdict, and confirm the regime by reading the pattern across several runs rather than a single one. The split is our labeled interpretation of how citation behavior clusters.

How many runs do I need before I call a query contested or settled?

More than most dashboards show you. SparkToro's 2026 study found less than a 1 in 100 chance that two runs return the same brand list, so a single run is one draw, not a measurement. The same study calls visibility measured as a rate across prompts run multiple times "a reasonable metric," so plan on repeated runs of every prompt on every engine, not a single check. Whatever your cadence, read the regime from the pattern across weeks of runs, per engine, and treat any change smaller than a few mentions as noise.

How do you win an unstable AI-citation query?

You do not win it outright; you earn your way in through a governed loop, and the engines decide. Measure the query's presence as a rate over weeks to confirm it is genuinely contested rather than just noisy. Diagnose it against the tells: low presence and a cited set that churns across runs. Execute consistent, fresh, human-approved content that gets your claim corroborated across several of the independent sources each engine samples, and make sure your own canonical page states the claim clearly, since ChatGPT has leaned toward official pages and documentation since August 2026. Then verify on re-runs of the same frozen prompt set, looking for a sustained climb rather than a one-run blip. The idea that agreement across sources settles the slot is our working hypothesis, not a proven result.

Free ChatGPT visibility check

See where AI answers skip your brand — then fix it, free

Start a free workspace with your domain: 10 buyer-intent prompts through ChatGPT show where you are mentioned, cited, or skipped, and who gets recommended instead. A free content audit in the same workspace hands you the first fix to ship.

Start Free

Free · No credit card · About 5 minutes

Related guides

Measurement

Why Your AI Visibility Score Moves

AI visibility scores move week to week even when nothing you control has changed. Here are the 6 sources of normal variance, the patterns that signal real movement, and what to watch.

Measurement

Most GEO Advice Is Untestable. Here's How to Run It, and Not Fool Yourself.

AI answers are non-deterministic, so a naive GEO test lies to you confidently. The four test-design rules that decide whether your AI-visibility result is signal or noise: measure a rate, test at the buyer's specificity level, test per language, and sort each cited source by the move.

How SolCrys Works

Don't delete 0-visibility prompts: a four-state lifecycle for AEO prompt sets

Most teams delete prompts that show 0 visibility. Eason Wang, SolCrys CPO, on why that's the wrong heuristic — a 0-visibility prompt with real buyer intent is a gap to fix, not a prompt to remove — and the four-state lifecycle (Keep + Act, Rewrite, Archive, Add) we ship to enforce it.