Source-led article
AI Visibility Rankings Unstable, Research Suggests Statistical Noise

New research highlights a significant challenge in tracking AI visibility: the instability of rankings. A paper by IQRush, with pre-release access provided to Search Engine Journal, suggests that generative AI models often produce varying responses, making citation shares and dashboard rankings mere snapshots rather than reliable, fixed data points. This variability implies that perceived differences between competitors might be due to statistical fluctuation rather than genuine performance gaps. This finding is corroborated by a separate team’s similar research published in April, underscoring the widespread nature of this issue.
The core problem stems from the inherent randomness built into generative AI models like SearchGPT, Gemini, or Perplexity. Repeatedly querying these platforms with the same question can yield different source citations each time. This means each citation is just one of many possible URLs the AI could have generated. Earlier work by the same author demonstrated this variability, showing that even with a clear lead, the margin of error could make it inaccurate to declare one source definitively outperforming another based on a single sample.
Por que importa
Key facts
| Aspect | Finding |
|---|---|
| Problem | AI visibility rankings are unstable and prone to statistical noise. |
| Cause | Generative AI models introduce randomness, leading to varied citation outputs. |
| Solution | Two-part stopping rule: ranking order must stabilize, and top sites must be clearly differentiated. |
| Implication | Marketers and SEOs need to approach AI visibility data with caution, avoiding single-measurement conclusions. |
The new paper aims to address a crucial question for digital marketers and SEO professionals in India: How much data is truly needed before AI visibility rankings become meaningful? The answer, according to the research, requires two conditions to be met simultaneously. First, the ranking order must stabilize, meaning new answers no longer significantly alter the hierarchy. Second, the top-ranked sites must be sufficiently differentiated, with differences exceeding the margin of error. If the top contenders are too close, the ranking is likely statistical noise rather than a reflection of true superiority.
In practical terms, this means that a simple increase in citation share after a content change might not necessarily indicate success. A three-point gain in SearchGPT citations, for instance, could easily fall within the natural variability between successive runs. To confidently claim a win, multiple measurements before and after any changes are essential. A single before-and-after comparison is insufficient to distinguish genuine impact from ordinary statistical noise. This is particularly relevant for Indian businesses investing in AI-driven content strategies and needing reliable metrics.
Contexto
The platform being measured also influences the amount of data required for reliable rankings. The paper notes that confidence levels are not uniform across all AI engines. For example, Gemini tends to concentrate citations on a few sites within a single answer, while SearchGPT distributes fewer citations across a wider range. This means each answer from SearchGPT carries more independent information. A budget that might provide sufficient confidence for Gemini rankings might leave significant uncertainty for SearchGPT, necessitating a tailored approach to data collection.
Sometimes, the most honest conclusion is that there isn’t enough data to make a definitive statement. The research found that three out of 30 tests failed to produce clearly separated top sites even after 125 questions. In such cases, the recommendation is to hold off on publishing a ranking that the data cannot support. A valuable tracking tool should be capable of indicating when data is insufficient, rather than always presenting a confident, potentially misleading, order.
For digital marketers and businesses in India relying on AI visibility tools, this research provides critical insights. It highlights the importance of scrutinizing the methodology behind tracking services. Rand Fishkin’s advice to ensure providers “show their math” becomes even more pertinent. Understanding if a tracker performs repeated checks and reports a range, or if it simply pulls data once and presents a clean, potentially deceptive, figure, is crucial for making informed decisions on AI content strategies.
Source: Search Engine Journal: https://www.searchenginejournal.com/ai-visibility-rankings-arent-stable-new-research-shows-its-mostly-statistical-noise/581905/