GEO
Back to Blog
Visibility MeasurementConsideration

AI Citation Volatility: Tracking the Same Prompt Set for Three Months

We tracked the same 40 questions every week across ChatGPT, Perplexity, and Gemini for three months. The cited brand lists overlapped by only about 55% on average, revealing how much AI citations drift and what makes a source more stable.

Tenten GEO TeamPublished 2026-01-075 min read
The same prompt flowing into several AI engines and producing shifting points of light that represent citation drift.

Ask ChatGPT the same question today and three months from now, and nearly half the brands in its answer may be different. That is not an illusion or evidence that the model has suddenly become worse. Generative AI citations drift, and the change is large enough to reshape a buyer's shortlist. We tracked one prompt set for a full quarter to see how large the movement was, where it occurred, and whether it followed any pattern. The result was clear: citation volatility is much greater than most brands assume, and stability reveals more than a one-time ranking.

What We Tracked and How

We selected 40 questions that B2B buyers genuinely ask, covering intents such as "best XX software," "XX tool comparison," and "which team is XX for?" Every week for three months, we repeated the exact wording in ChatGPT, Perplexity, and Gemini. Each run recorded the brands in the answer, the cited URLs, and which brand appeared first.

Controlling the variables was essential. We kept the same account settings and region, and disabled personalized memory and search history so changes were more likely to come from the models and sources rather than our own activity. The prompt wording never changed. Once a word changes, the system may retrieve different material, turning the run into a new question instead of another point in the same time series.

How Much Changed After Three Months?

Volatility did not appear only in obscure queries. Even for a crowded topic such as "best project management software," nearly 40 percent of the cited brand list changed over three months. A position that appears secure in one report may already be loosening.

  • Brand lists at the beginning and end of a month overlapped by about 55% on average, meaning nearly half of the cited set could change within a quarter.
  • For roughly one-third of the questions, the first brand mentioned changed at least once during the three-month period.
  • Perplexity moved the most because it relies heavily on live search. ChatGPT was more stable for established authorities but still volatile for emerging brands.
  • Long-tail and niche questions reshuffled their brand lists roughly two to three times faster than popular, broad questions.
  • New content with clear structure, definitions, and data could enter the citation set within two to three weeks of publication.
Citation stability reflects a brand's real position better than a one-time ranking. A single article can create a short-lived spike, but repeated citations show that models continue to treat the brand as a trustworthy source.Tenten GEO Brand Radar tracking
Line chart showing cited brand lists changing over three months, with an overlap rate of about 50%.
When the same questions were repeated weekly for three months, cited brand lists overlapped by only about 55% on average.

Why Does the Same Question Produce Different Answers?

Several layers of change compound one another. First comes retrieval: a model searches the web or its index before answering, while search results are reordered daily as new pages appear and old pages fall away. Next come model updates, with providers adjusting weights and safety behavior every few weeks. Finally, generation itself contains randomness, so wording and examples can move even when the first two layers remain unchanged.

Competitors create another source of movement. When one publishes a well-structured comparison with precise definitions, it can replace an existing citation within two or three weeks. An AI answer is a balance that is continually recalculated, not a permanent shortlist.

The Movement Still Follows Patterns

Some of the apparent disorder is predictable. Authoritative brands supported repeatedly by third parties moved less. Brands with thin information and only self-published claims dropped out most easily. Questions with clear intent and widely accepted answers were more stable, while vague, subjective questions with many possible choices remained volatile. Better content structure and independent corroboration can therefore move a brand toward the stable end of the range.

How Brands Should Respond to Citation Volatility

The least useful response is to screenshot AI answers every day and rewrite an article whenever the brand disappears. Daily lists are noisy by nature, and decisions based on one data point waste time. Extend the measurement period and use trend indicators such as citation share and list overlap to determine whether the brand is becoming more stable or losing ground.

Improving stability requires clear definitions, numbers, and context that models can extract without guesswork. It also requires consistent brand and entity signals across credible sources, plus independent mentions that give AI systems more than one reason to trust the claim. Brand Radar is designed for that work: it monitors citation movement across generative engines continuously and turns daily noise into a trend a team can use.

AI answers will not settle into place for you. The durable strategy is to become a source that is difficult to replace. Book a 30-minute GEO diagnostic session to measure how stable your citations are, identify the questions where they fail, and review the pattern with your own brand data.

Frequently asked questions

Can an AI-cited brand list really change that much within a few months?
Yes. When we repeated the same questions weekly for three months, brand lists at the beginning and end of each month overlapped by only about 55 percent on average. Nearly half the cited set could change, even for popular topics.
Why does the same AI question return different brands?
Several moving layers affect the answer: live search results reorder daily, models are updated every few weeks, generation contains randomness, and competitors keep publishing new material. The result is continuously recalculated rather than fixed.
How should a brand track citation volatility?
Do not rely on a single day's ranking. Track citation share and list overlap over a longer period to see whether the brand is becoming more stable or losing ground. Treat the two to three weeks after publishing new content as the key observation window.

READY WHEN YOU ARE

How visible is your brand in AI answers?

In a 30-minute GEO diagnostic session, we use real prompts to identify your visibility gaps across major AI engines and show you what to fix first.

Book a 30-minute diagnostic