Why ChatGPT's Citations Keep Changing: The Four Hidden Search Pipelines

ChatGPT routes queries through four hidden retrieval systems, new research shows, which explains why the same question can produce different brand citations.

ChatGPT does not run one search system behind its answers. Independent analyses of its raw network traffic found four hidden retrieval pipelines, internally labelled Labrador, Bright, Oxylabs and SERP, each pulling from a different pool of sources. When ChatGPT switches between them, the citations it shows can change even though your content hasn't.

What changed?

Two SEO researchers inspected the raw network traffic ChatGPT sends behind its answers, rather than relying on the polished citation cards users see. According to Search Engine Land, they found internal source-selection labels sitting behind the answer: Labrador, Bright, Oxylabs and SERP.

Chris Green ran the larger test. He tested 1,000 prompts up to 10 times each and captured 9,946 completed search runs. Most prompts stayed on one retrieval source: Labrador accounted for 88.1% of primary search sources, Bright for 9.9%, Oxylabs for 1.7%, and SERP for 0.3%. But 11.6% of prompts changed primary search source across repeated runs.

The consequence of that switching is the real story. URL overlap dropped from 0.273 to 0.149 when the search source changed, and domain overlap fell from 0.265 to 0.155, roughly 45% lower URL overlap and 42% lower domain overlap. In plain terms, when ChatGPT switches pipelines, nearly half the sources it would have cited disappear from the answer, replaced by different ones.

Suganthan Mohanadasan, co-founder of Snippet Digital, ran a separate analysis of his own account's traffic. Writing on Search Engine Journal, he found the labels point to Bright Data and Oxylabs, two commercial scraping firms that are direct rivals, handling ChatGPT's open web fetching. Bright did the bulk of the fetching, especially on commercial, shopping, finance and weather queries, while Oxylabs skewed regional and local, Labrador stayed on news and reference, and SERP mostly turned up on news.

He also found that not every query reaches these pipelines at all. ChatGPT classifies some queries with a field called turn_use_case before deciding whether to search, meaning some prompts are filed as text and skip web search entirely, even when they sound current. That routing decision determines which pages ever get a chance to be read, let alone cited.

Mohanadasan's traffic also separated three outcomes that are often lumped together in GEO advice: fetched, cited and mentioned. A page can be fetched into ChatGPT's context without being shown to users. It can be cited as the source behind a specific sentence. Or a brand can simply be mentioned without being the source of the claim. In his sample, Reddit and YouTube were both fetched often, but Reddit was cited and YouTube was not, which he attributed to text availability: Reddit threads expose text, while YouTube search results often provide metadata rather than transcripts.

Both researchers are careful about the limits of their samples. As Search Engine Journal notes, this is one person, one logged-in Pro account, a few days of traffic, not a population study, logging around 1,240 source records across a few dozen searches. The structural finding, that these pipelines exist and behave differently, holds up. The exact percentages are a snapshot, not a fixed rule.

Why does this matter for your business?

If you've tracked your brand's ChatGPT citations and watched them swing week to week with no obvious cause, this is likely part of the explanation. You're not being tracked by one ChatGPT. You're being tracked by several retrieval systems that share an interface but pull from different pools of sources.

This has direct implications for pricing pages, product specs and comparison content. Mohanadasan's traffic showed vendor pages were cited for their own facts, such as prices and specs, while third-party pages were more likely to support broader recommendation claims. That splits your visibility strategy into two jobs: making your own site machine-readable for facts, and making sure independent sites carry the opinion-based claims that win recommendations.

Readability matters more than most brands assume. Both analyses showed ChatGPT's source selection depended partly on what it could actually retrieve and read. Mohanadasan found cases where ChatGPT appeared to prefer official pricing pages, then fell back to third-party sources when prices were hidden behind JavaScript or otherwise hard to parse. A price buried in a JavaScript widget isn't just a UX flaw. It can be the reason a competitor's third-party review gets cited instead of your own product page.

This lands at a moment when ChatGPT's referral traffic is climbing fast. Weekly analysis from Anicca reports that SE Ranking analysed traffic data from 101,574 websites across 250 countries and territories and found ChatGPT's share of referral traffic rose from 0.23% in April to 0.32% in May, its highest level on record and 8% above the previous peak from October 2025. More traffic is riding on a citation system that is proving less predictable than most brands assumed.

What should you do now?

Stop treating "ranking in ChatGPT" as a single target. Multiple pipelines with different source pools mean consistency, not a one-off fix, is the real challenge.

Prioritise plain, crawlable HTML over content locked behind scripts. The research is consistent here: plain HTML, crawlable facts, clear pricing and specs, strong third-party coverage, and text-heavy pages all became more important once source selection depended on retrieval and readability.

Build presence on the third-party sites that actually get cited for opinions and recommendations, not just your own domain. Track outcomes separately too, because being fetched, mentioned and cited are three different things, and only the last one drives the specific sentence a customer reads.

Because AI Overviews, ChatGPT and other assistants change their retrieval behaviour from week to week, it's worth checking where your brand currently stands with a free audit at Sited rather than assuming last month's snapshot still holds.

Finally, treat any single vendor's percentages as directional rather than definitive. As the researchers themselves caution, the numbers move even when the structure holds, so re-check your visibility regularly rather than optimising once and walking away.

Frequently asked questions

What are Labrador, Bright, Oxylabs and SERP?

They are internal labels found in ChatGPT's network traffic marking which retrieval provider fetched a given web result. A result_source field is attached to web results, with Labrador covering established publishers and reference sites, Bright tied to Bright Data, Oxylabs tied to Oxylabs, and SERP an open-web baseline that appeared mostly in news-style results.

Does this mean my SEO work is wasted?

No, but it means SEO alone isn't sufficient. Readable, well-structured pages still need to win the fetch-and-read step before they can even be considered for citation, and third-party coverage remains critical for recommendation-type claims.

Why do my ChatGPT citations change from week to week?

Partly because your query may be routed to a different retrieval pipeline than last time, each with a different pool of sources. Green's research found that URL overlap dropped from 0.273 to 0.149 when the search source changed, and domain overlap fell from 0.265 to 0.155 between runs.

Is this the same for every ChatGPT user?

Not necessarily. Both studies were based on limited, individual-account samples, and the researchers themselves warn against treating specific percentages as universal, since the numbers come from a relatively small set of queries on single accounts.

Should brands try to target a specific pipeline like Bright or Oxylabs?

You can't choose which pipeline handles a query, so the more useful approach is making your content easy for any scraper to read: plain HTML, visible pricing, clear specs, and strong coverage on third-party sites that these systems already trust.

Sources