Reverse engineering · July 2026

What ChatGPT pulls,
what it shows, what it cites.

27,000 pages put under the microscope. The engines behind the answers, the two hundred characters that stand for you, and the caches that freeze them.

URLs retrieved by ChatGPT58,000
Promoted to sources7,600
Cited in the answer text5,000
Opened and read by the model760
Powered by RESONEO

When ChatGPT answers using the web, it does not go and read your page. It queries one or more engines, takes away a title, an address and roughly two hundred characters, and works from that. The page itself is opened in barely more than one case out of 80, and in thinking mode only. The rest of the time, what stands for you in front of the model fits in three lines of text produced elsewhere, at another moment, and sometimes served to someone else before you.

This study takes that chain apart layer by layer, across more than twelve hundred real answers captured in July 2026. Who brings the results back, what the model sees of them exactly, how long that copy stays valid, and what clears the final step to become a visible link in the answer.

Table of contents

1 On July 21, OpenAI turned off the light 2 The real ChatGPT is the free one 3 What the model sees of you 4 What ChatGPT finds, what ChatGPT reads 5 Cited, displayed, invisible 6 What gets opened gets cited 7 Five verticals, five separate systems 8 What we still do not know 9 Odds and ends we also found 10 Method: how we know 11 Credits, sources and tools Crossed findings with Jérôme Salomon (Oncrawl)
1

On July 21, OpenAI turned off the light

In the code, every result used to carry the engine that brought it back, through the result_source field. That field has vanished from the stream. We are working on a window that has just closed.

Before, and after measured 26/07/2026
Before 21 July
Every result names its engine. Google Shopping tokens are readable in the clear inside the carousels.
On 21 July
The field naming the engine disappears from the stream. In its wake, Google Shopping tokens leave the product carousels: the query and provider fields now contain only the word deprecated.
After
No answer says where its results come from any more. From now on you recognise an engine by the shape of what it delivers.
Evidence

The result_source field was carried by every result object in the server stream, with four possible values: labrador, bright, oxylabs, serp. Across the 1,249 answers in the corpus, 760 carry it and 489 do not. Last labelled stream on 21 July, first fully bare day on 22 July.

OpenAI is trying to shed Google but cannot yet format signature

Two of the four engines, bright and oxylabs, serve scraped Google. We established it by replaying the same queries on both engines, then looking for the URLs used by ChatGPT in their results. One URL in three shows up on Google's first page. Only one in twenty on Bing. And when it is in Google, nine titles out of ten match character for character, ellipsis included, in the same spot.

Google
1 in 3
of bright URLs found on the first page
Bing
1 in 20
of the same URLs
labrador
outside both
no engine fingerprint

What makes the gap conclusive is that Google and Bing barely overlap with each other on these queries. Being present in one and not in the other therefore becomes a strong signal rather than a ranking coincidence. As for labrador, it belongs to neither: it is an in-house index, not a scrape.

In thinking mode, ChatGPT runs three quarters of the time on bright. So on Google.

The same move shows up on the shopping side. Google tokens have left the product carousels, but the Google product id survives in the clear inside the request fired when a user clicks a card, this time alongside an encrypted block.

Google has not left shopping. It has been encrypted.

And the dependency is not residual, it is functional. The offers displayed in the carousels do come from OpenAI's internal catalogue, fed by merchant feeds. But that catalogue cannot do everything: Google remains the oracle on best price, on part of the product metadata and on reviews. In other words, OpenAI has made its dependency unreadable in the visible stream, while keeping it operational in the layer that only fires on click.

A date match?

The removal falls two days after Google's complaint against SerpApi was dismissed. We flag it as a calendar match, not as a cause-and-effect mechanism. Nothing in our measurements links the two, but you have to admit it is unsettling ;)

2

The real ChatGPT is the free one

The free tier does not reason. Across our 493 free-tier answers, not a single one went into thinking mode. And instant settles for a single fan-out in more than eight cases out of ten, and never opens a page to read it.

Two regimes, two orders of magnitude measured share
Free, instant mode
Paid, thinking mode
Number of search fan-outs
1
6
Pages pulled
11
100
Distinct domains
8
28
Pages actually opened
0
1 conversation in 6
Dominant engine
labrador
scraped Google (bright)

Median values per answer. Thinking mode reaches nine times wider, and sometimes up to several hundred pages on a single question.

Your question picks the engine

In instant mode, on a free account, four families of questions and only two behaviours.

What you ask for labrador Google No snippet
Stable information100 %0 %1 %
A business, a place100 %0 %25 %
A product to buy100 %0 %1 %
Live information50 %50 %8 %

The last row is the questions whose answer changes within the day: who won the game last night, the Masters leaderboard live, today's Fed decision. They are the only ones that open the door to scraped Google. Everything else, and that is most of what people ask ChatGPT, never leaves the in-house index.

In other words, understanding labrador is not a niche topic: it supplies almost every result, for the population that actually uses ChatGPT.

Exact values

Search results only, empty snippets excluded from the share, across 5,245 snippets and 424 free answers in instant mode. Stable information: 99.9 % labrador across 2,136 snippets. Business and place: 100 % across 310. Product: 97.6 % across 1,392. News: 52.6 % labrador and 46.1 % Google across 1,407. The split is read off the length of the served snippet, which separates the two engines unambiguously: past one hundred and ninety-five characters the source is labrador in 99.9 % of cases, measured over the period when the stream still named its supplier.

The missing-snippet column deserves to be read alongside the first: on businesses and places, a quarter of results arrive with no snippet at all, just a title and an address. The best-served class in the table is also the one where the model receives the least material.

In thinking, everything flips

The table on the left is for instant mode. Ask the same questions again in thinking and the ratio turns over, across all four families at once.

The same question, in thinking labrador Google
Stable information1 %99 %
A business, a place3 %95 %
A product to buy1 %99 %
Live information0 %96 %

In instant, the family of the question decides everything. In thinking, it has no effect at all: the four values close up where instant spreads them from 50 to 100 %. Businesses and places even change architecture between the two regimes. In instant, the answer is carried by a map, and the result list is just a reminder of links, a quarter of which have no snippet at all. In thinking, the map disappears entirely, the list goes back to being a real retrieval net, and the share of snippet-less results drops from 25 % to 4 %.

Labrador is neither rebadged Bing nor a licensing catalogue

Two readings of labrador circulate, and our measurements rule out both. The first makes it Bing served under another name, on the grounds of the OpenAI-Microsoft partnership. One measurement is enough to overturn it: Bing cuts the titles it displays at seventy-five characters, never one more, and a quarter of labrador titles run past that ceiling. The longest one we have seen runs to nine hundred and seventy-seven.

Labrador is not rebadged Bing. A quarter of its titles are too long to be.
A piece of information even SEOs tend to miss

The maximum length of title tags displayed in Bing's result pages is seventy-five characters, never one more. That ceiling is what ruled out the hypothesis that labrador is rebadged Bing: a quarter of labrador titles cross it, which is physically impossible for a Bing scrape.

A fifth engine, named bing, has been reported by other researchers. We have never observed it, on any account, in any regime, on any date. Our hypothesis is that it may be exposed only to OpenAI Teams accounts, because of the Microsoft relationship.

Jérôme Salomon has passed us June captures where that pipeline shows up in bulk on a Team account, while the matching free account sees none of it on the same day (collaboration section).

Two other measurements point the same way. Snippet lengths are incompatible: Bing cuts wide and irregular, labrador stops dead at two hundred and two characters, across every one of its results. And the labrador snippet is no engine's snippet: on labrador URLs that also rank on Google and on Bing, text reuse is nil on both sides. The same measurement applied to bright finds Google's text four times out of ten.

The second reading makes labrador the preserve of publishers under agreement with OpenAI. Our measurements do not support that one either. At equal engine, having a licensing deal or not changes nothing about the treatment: hundreds of outlets with no agreement get exactly the same format, the same length, the same freshness as the partners.

Labrador is better described as a repository of immediate access.

Labrador = an in-house index, topped up with press feeds, open scientific repositories and partner platforms.

What defines it is not the contract, it is that OpenAI queries it without paying a third party and without waiting for a scrape.

Evidence

All engine shares in this section are computed on results of type search only, across the 373 answers before the cut-off, and read straight from the result_source field, with no classifier.

This filter is essential. Counted across all types, oxylabs looks like 12.7 % of instant mode. On search results alone it drops to zero: oxylabs is a news pipeline and nothing else, all of its results are of type news. Symmetrically, scientific results and forums are served 100 % by labrador, so including them mechanically inflates its share.

Exact values, search results only. Free and instant: labrador 63.6 %, bright 34.7 %, serp 1.6 %, across 1,793 results and 158 answers. Paid and thinking: bright 74.6 %, labrador 23.6 %, serp 1.9 %, across 16,407 results and 176 answers.

Both averages hold, but they depend entirely on what is being asked. The batch that produced them was half made of news questions, the one family that goes to Google. On a batch of stable questions, labrador's share on the free tier rises to 100 %. An engine average therefore only means something alongside the mix of questions that produced it: it is the table by family, above, that describes how the system behaves.

3

What the model sees of you

In instant mode, everything fits in a title and two hundred characters. And those two hundred characters are not the ones you think.

A reminder before going into detail

Once its fan-outs are fired, ChatGPT receives a list of results exactly like the one you get when typing a query into a search engine: for each page, a title, a snippet and a URL. Nothing more. It is on that grounding net, and on it alone, that the model relies to write its answer and decide who it will cite.

In other words, those three fields are your only representation in front of the model in the vast majority of cases. Hence the importance of knowing exactly how they are built.

The title: whole, or cut at sixty characters format signature

Depending on the engine that brings it back, the same page arrives with its full title or with a title already truncated by a results page. Through labrador, the tag passes whole: a quarter of titles run past seventy-five characters and almost none end in an ellipsis. Through scraped Google, the title is cut to an engine's display length, and nearly one in four carries the mark of that cut.

labrador
Titles over 75 characters24 %
Truncated titlesnear zero
scraped Google (bright)
Titles over 75 characters1.5 %
Truncated titles23 %
What a result looks like, depending on the pipeline that brings it

Here are the three formats actually passed to the model, from shortest to longest. The same structure every time, a URL, a title and a snippet, but budgets that vary sevenfold.

Result served by the bright pipeline
bright, the Google scrape. Title of 48 characters at the median, cut off at 60, with an ellipsis one time in four. Snippet of 159 characters, in search-results format.
Result served by the labrador pipeline
labrador, the in-house index. The title passes whole, never truncated, 58 characters at the median and up to several hundred. The snippet is anchored on the page H1 and cut at 202 characters.
News result served by labrador
labrador, news. Same pipe, entirely different regime: around 1,100 characters at the median, seven times the previous format, and a clean cap at two hundred words. It is not an extract from the page but a rewritten summary, impossible to find verbatim anywhere.
The labrador snippet: two hundred characters taken from the start of your page body format signature

The snippet stored by labrador's in-house index is cut off just past two hundred characters, and it is taken from the start of the page's rendered body, not from your metadata. Your meta description is ignored. It stays useful for bright results, the Google scrape, which picks it up one time in three.

When a page has an H1, the snippet contains it eight times out of ten.

The cut almost never falls in the middle of the H1: across nearly four hundred snippets containing it, only four truncate it. Your grounding budget is therefore perfectly predictable. It is made of three blocks that always follow the same order: whatever sits before the H1, the H1 itself, and whatever room is left afterwards for your actual content.

7
51 chars of H1
146 characters to convince
7 chars before the H1, on average 51 chars of median H1 146 chars of actual content

What eats the start of the snippet is not what people assume. Breadcrumbs and bylines are marginal. The two real losses are the section kicker and the alt text of the first image, the one sitting right after the title, which on its own can eat fifty characters.

What comes before the H1FrequencyCharacters lost
Kicker, section, category29 %18
Publication or update date11 %25
Alt text of the first image, right under the title9 %50
Site or brand name7 %44
Breadcrumb7 %23
Author byline3 %45

The first lever is not there, though. Across the pages we examined, one in seven simply has no H1 markup at all. In that case the anchoring becomes unpredictable: the snippet starts on some subheading picked by the template, and you no longer control anything.

Evidence

Method: 534 pages actually cited by ChatGPT were fetched then compared with the snippet the index had stored for them. Of the 463 that have an H1, 387 snippets contain it, that is 83.6 %. Among those 387, the H1 is the very first character in 297 cases, that is 77 %. The prefix, when there is one, is 20 characters at the median.

The snippet is not chosen based on the question asked: seven regional variants of the same press release, queried with the same request, produce seven different anchor points. And it is not built at search time: the vast majority of pages seen under several different queries return a strictly identical snippet. It is frozen at crawl time, and it ages.

What to do with this. Your grounding budget in instant mode is your full title plus roughly one hundred and fifty useful characters starting at your H1. Mark up an H1, clear the runway between it and the first paragraph, and keep an eye on the alt text of the first image, the one sitting right after the title.

4

What ChatGPT finds, what ChatGPT reads

OpenAI keeps your pages in two distinct stores, written by two different robots. The copy you have in there has an age. You do not choose it, and you do not see it in your logs.

The index and the read cache

Everything else in this section reads through this one distinction. The index is what finds you, the read cache is what reads you. Not the same content, not the same robots, not the same delays.

The index The read cache
What it holds a short snippet, frozen at crawl time the whole page, already converted to Markdown
You reach it through a query, some keywords an address, the exact URL
What it is for finding your page reading it
Who writes it OAI-SearchBot ChatGPT-User, and OAI-SearchBot too
What updates it a SearchBot visit, and nothing else a user request, once past thirty minutes
Lifetime no expiry observed; 13 % of snippets run more than a month behind the crawl no expiry observed; several weeks, if not months
How to measure it the API Crawled field, which leaves no trace in your logs ChatGPT-User showing up in your server logs

The two do not hold the same pages. Jérôme found that not every URL in the index turns up in the read cache. Being findable and being readable are not the same thing.

How the read cache ages format signature
Freshness
30 minutes

Below that window, the page is not reopened at all. One user asks for your page at 2 pm, another asks for it at 2:29 pm: the 2 pm copy is what the second one gets, and ChatGPT does not even bother going back to check whether the page changed.

Retention
Several months

With no new request, the copy is not discarded. We observed no expiry.

The mechanism is a cache that serves first and checks afterwards. Once the copy is past the half hour, ChatGPT still shows you the old version, then goes off to refresh in the background. The fresh version will only be served on the next request. In other words, when you see ChatGPT-User in your logs, it is not a read for the current user: it is a refresh of OpenAI's cache, which the next visitor will benefit from.

The chain of evidence

On our test site, a random fingerprint is written at the same instant into the served page and into our server logs. When the model spits it back, we know exactly which visit produced the copy it just read.

Here is the sequence that established the cache is shared between users. A fresh page, never read by anyone, and two paid accounts with no connection between them.

1. Account A asks for the page our log: one ChatGPT-User hit
fingerprint written = 2a2b

The copy has just been made.

22 min later →
2. Account B, other session, other country model answer:
« ... 2a2b ... »

It hands back, word for word, account A's fingerprint.

3. What our server saw in between no hit at all

Account B never touched our server. It read account A's copy.

Double proof: the fingerprint is identical, and the hit counter stays at zero. Either one alone would prove nothing.

1
The cache is shared between users

The copy created by one account is served back to another account, in another country, on another plan, without a single hit reaching the origin server. The cache key is the page address, not the user.

Another point worth knowing: our test pages sent a Cache-Control: no-store header, which explicitly forbids any caching. They get cached anyway.

2
Three robots, one index and one cache

ChatGPT-User builds the read cache, the whole-page one: it is the robot that fires when the model or the user explicitly asks for a page. OAI-SearchBot, the search robot, builds the index, the snippet one, and it also feeds the read cache, though not systematically. Pages it visited alone, with no on-demand read ever fired, are indeed in there.

3
ChatGPT-User never updates the snippet

The reverse is not true, and that is the counter-intuitive part. ChatGPT-User never touches the index. Across pages dated both by OpenAI's index and by server logs, the served snippet tracks the SearchBot visit within a few days; yet ChatGPT-User is the more recent of the two robots in most cases. So it does come back to read your pages, it does refresh its whole-page copy, and none of that changes the snippet search displays.

4
The training robot gets into neither

Almost every index entry is more recent than GPTBot's last visit, so refreshed after it came and not by it, and we had verified the same thing on the read cache. The practical consequence is worth stating: being crawled and being available are two different things. The detail of that measurement sits in the collaboration section.

The peripheral caches

Beyond the two main stores, ChatGPT also keeps objects it built itself. They are not read like pages, but they age the same way.

What is cachedObserved lifetimeConsequence
The summary of a news articlestable across capturesGenerated once, served repeatedly afterwards
The product or entity card (ChatGPT sidebar)weeks or monthsIdentical from one account to the next, across countries too

On the product card, the demonstration is striking: the same product clicked from a French paid account and from an American free account returns exactly the same text, generated once.

Evidence

We bracketed the freshness window by dichotomy on two different pages and two different days, by reading the served fingerprint while checking in parallel for the presence or absence of a server hit. No refresh at 3, 16 and 25 minutes. Refresh at 31 and 33 minutes, then at 1 hour, 9 hours and 10 days. The threshold therefore sits between 25 and 31 minutes, and 30 is a plausible round value rather than a value we read anywhere.

Retention is a lower bound: a copy from 11 July was still being served on the 23rd. We never observed an expiry, which does not prove there is none. On Jérôme Salomon's testing ground, pages were still being served from this cache more than ninety days after they were fetched, with a server log to back it.

Collaboration

Crossed findings with Jérôme Salomon

Two investigations run in parallel on the same questions, one at Oncrawl, one here. Here is what came out of putting them side by side.

Jérôme Salomon

Throughout these investigations we worked in dialogue with Jérôme Salomon, technical SEO expert at Oncrawl. Each side ran its own measurements, with its own tools and its own testing ground; at every exchange, one set of findings corrected or extended the other. Here is what that collaboration established, between cross-confirmations and complementary finds.

The cache dossier cross-measured, July 2026

The previous section owes almost everything to this part of the dialogue. Two independent protocols answer each other here: server logs and the API on one side, watermarked pages and the account fleet on the other. Five facts, established by Jérôme.

  1. A parameter queries the cache with no web access. Passing external_web_access: false to the API asks the model to answer without going out to the web. Whatever it hands back then necessarily comes from its store. That is concrete proof that the discovery index and the read cache exist, and the starting point of the whole previous section.
  2. The index is dated by OAI-SearchBot, and the two stores do not hold the same pages. On one section of Oncrawl's site, dozens of pages were dated both by OpenAI's index and by the server logs: the served snippet tracks the SearchBot visit within a few days. That robot feeds both stores, since pages it visited alone also turn up in the read cache, about half the sample. They do not hold the same pages, though: pages clearly present in the index were absent from the read cache, which confirms that search serves from the index. So the index is not empty, it is selective: fed by SearchBot and by partner feeds. One caveat: refresh-on-SearchBot-visit is the rule, not a law. A minority of snippets run months behind, and a snippet can keep ageing through a re-crawl.
  3. The noindex directive is ignored. The analyses Jérôme ran confirm it: pages carrying meta robots: noindex and visited by OpenAI's robots still end up in the cache. Forbidding indexing does not keep you out of the store.
  4. Being read does not guarantee being stored. A ChatGPT-User visit is no guarantee of making it into the cache, at least not in what we are able to observe. Of the pages that robot actually visited, at least half do turn up there. We have no proof that all of them do, and the limit comes from our instrument: the model is the one deciding whether to open a page, so a test produces a false negative as soon as it does not attempt the open. That half is a floor, and the robot's visit probably triggers the caching mechanically. The hack described below now makes it possible to redo the measurement without asking the model anything.
  5. A retention dated at more than ninety days. Pages whose last logged OpenAI robot visit went back ninety-one days were still being served from the cache. Unlike the other retention measurements, the starting point is not an estimate: it is a line in a server log, not an inference. Not to be confused with the thirty-minute freshness window from the previous section. This is how long a copy survives, not how long before it gets rechecked.
The web_search API as a measuring instrument not contractual

Jérôme's hunch was right: the API exposes the same machinery as the application, down to its internal markers. The engine itself cannot be picked from the API, but the lead pointed at the right place: digging into it turned up the field nobody looks at.

The lever fits on one line. On POST /v1/responses, with the web_search tool, you ask for include: ["web_search_call.results"]. Every result then arrives prefixed with internal metadata the application stream never exposes.

[wordlim: 200] Published: today; Crawled: 6 days ago; <page content>

Three pieces of information, across the 7,782 results we analysed through the API. Crawled gives the age of the copy being served, URL by URL, with no access to the site's logs whatsoever: an audit oracle that exists nowhere else. wordlim gives the word budget allocated to the domain. And the citation marker names the source channel: search, news, science, forum, opened page. None of it is documented, and nothing stops OpenAI from removing it tomorrow, the way the field naming the pipelines was removed on 21 July.

The map of word budgets is OpenAI's publisher contracts, measured from the outside.
Budget per source Who it applies to Reading
25 words Guardian, ESPN, Washington Post, Bloomberg, Corriere, Spiegel… Cause undetermined
100 words 11 domains identified, all of them known licensees: Le Monde, WSJ, Politico, Bild, Condé Nast, Hearst… Looks like a contract clause
200 words The default, all the rest of the web The normal regime

The value is a constant per domain-and-channel pair, with no residual variance at all once the right key is found, and a domain's PDFs inherit the domain's budget. The 100 class explains itself. The 25 class stays a puzzle: it mixes publishers with an agreement and publishers without one, and the robots.txt explanation is refuted by counter-examples in both directions.

The instruction was spelled out in the system prompt

That word cap is not an API invention. It sat in the ChatGPT system prompt we published in March: the [wordlim N] instruction, with two hundred words per source by default, is present in the GPT-5.3 version. The API now confirms it server-side, source by source. The loop closes between what the prompt ordered yesterday and what the stream measures today.

### Copyright and word limits - If you derived any information from a webpage source, you MUST cite it. Any part of your response that used information from sources must have citations. Do NOT miss any citations, otherwise it would result in copyright violations. - You must cite all the trustworthy sources that support a claim or statement in one cite block, and order them by how well they support the point. - Quotes: ≤10 words for lyrics; ≤25 words from any single non-lyrical source. - Per-source paraphrase cap: respect `[wordlim N]` (default 200 words/source). Do not exceed; caps add across cited sources. - Don't reproduce full articles/long passages; use brief quotes + paraphrase/summaries. - Exception: these quote/paraphrase caps do not apply to reddit.com.

Excerpt from the system prompt, published in our March study.

The age oracle, for its part, holds two surprises for the general reader. On the live web it does date a real read, decoupled from the publication date. But on corpora ingested in bulk, it dates the content of the copy, and that copy was never refreshed.

Wikipedia runs two months late

A Wikipedia page we analysed, edited several times a day, is served with a "Crawled: 2 months ago" marker. Across 308 Wikipedia results, 209 are dated between eight and ninety days: the encyclopedia models cite most is served from a periodic snapshot, not from the site. Those ages fit the encyclopedia's own rhythm, whose full exports are produced once a month: the copy follows the dump schedule, not the edit schedule.

Reddit is frozen at thread creation

On Reddit, the collection date is the date the thread was created, zero days apart at the median, going back as far as seven years. A thread opened in 2019 is served as it stood at creation, or a few weeks later: replies added after that are not in the copy.

News publishers, seen from both sides

This is the ground where our two observation posts differ the most, and overlap the best. Three points.

  1. Reuters, or the proof of a channel that runs on zero crawl. Reuters is the first news domain served in the application, with 120 occurrences in our corpus. In the API it cannot be found: ten dedicated calls, zero sources, zero opened pages, the model answering that it cannot reach it. The hole is surgical, every Reuters satellite domain comes out normally. And the demand does not vanish: the model spontaneously fires 47 site:reuters.com queries, all of them empty. So there is a direct ingestion channel for press content, one that does not depend on the crawl. Jérôme supplies the mechanism: partners have an API to push their articles to OpenAI, and those articles are available without ever having needed a first crawl. They do get crawled afterwards, by ChatGPT-User and by SearchBot, just less than the articles that do not go through that channel. The concrete consequence for publishers: the best integrated are also the least crawled, and no ChatGPT in your logs does not mean no ChatGPT in the answers.
  2. News snippets are machine-produced summaries. Around eleven hundred characters at the median, capped at two hundred words, with a hundred-word step on a few hosts, and never findable verbatim on the web. We measured it on the application stream; Jérôme's knowledge of that sector supports that reading. It also gives a way to read an article's route in: pushed through the partner API, it arrives with the news flag; picked up by the crawl, it arrives with the search flag. A 30 July capture shows both sides in the same answer, for the same publisher: Le Monde appears seven times as news with a 180-word summary, and once as search with a 32-word snippet. The flag does not describe the topic, it describes the door.
  3. Bing does exist, and the best lead remains Team accounts. In June, on a Team account, Jérôme observed the bing pipeline in bulk, on 93.7 % of the sources brought back. That same month, on our accounts, zero. We ruled out two explanations: it is neither the server region that handles the answer, nor free versus paid. What remains is the Team plan: the most likely hypothesis. We suspect an arrangement between OpenAI and Microsoft opening Bing access to ChatGPT Team users, without being in a position to verify it.
5

Listed, promoted, cited

Everything ChatGPT pulls ends up listed somewhere. But between being listed and being cited there are three steps, and almost nothing climbs them.

Nothing in the data marks a result as retained. The notion of a citation only exists at display time, which is why published figures on the subject diverge so much from one study to the next: each counts at a different layer without saying so. We rebuilt the layers from the raw stream, then checked how they map onto the interface across two real sessions where the client names each display zone itself.

A correction we owe to that check

It is tempting to write that most pulled pages are never seen. That is false, and we believed it for a while. Every page brought back by retrieval is indeed listed in the panel that opens when you click "Sources" under the answer. Across more than sixty thousand answer-and-page pairs, not one is absent from every display surface.

What is true is subtler: nine out of ten stay at the bottom of that panel, in the "More" section, behind one click, never promoted to the top and never picked up in the text. Technically visible, practically invisible.

Being cited is not binary

A citation anchor in the text can carry several sources. The pill shows the name of the first one and flags the others with a "+1", reachable by opening the tooltip. All of them are attached to the citation, but only one gives it its name.

That is the whole gap between the two steps below: roughly 1/3 of URLs are attached to a citation without being its lead source. They exist, they are clickable, but they are not what the reader sees first.

The first column follows the Sources panel, step by step. The second column says what the model actually did with the pages. The filters cross the mode, the nature of the result and the engine; combinations with too little data are flagged.

The side panel, step by step
Every URL in the side panel61 332
Attached to a citation tooltip 7 616
First citation visible in the tooltip 5 032
ChatGPT citation tooltip showing 1/2
The pill carries the name of the first source and a "+1". The tooltip reads "1/2": a second source is attached to the same citation anchor, reachable with the arrows.
What the model did with them
759
pages opened and read by the model, in thinking mode only
53 716
stay at the bottom of the panel, under "More", never promoted nor cited
The vast majority of pages are cited without ever having been opened by the model. Opening remains the exception.
Per answer, the order of magnitude

A median answer lists around twenty pages at the bottom of the sources side panel, promotes five to the top, cites three in its text and opens none. The median for opened pages is zero: nine answers out of ten never open anything.

Nine pages out of ten never leave the bottom of the panel. They were selected, passed to the model, they potentially contributed to the answer, but ChatGPT gives them almost no prominence.

Two rival systems, the same idea

ChatGPT encodes its citations as turn0search11: turn zero, search family, eleventh result. Gemini encodes its own as PerQueryResult(index='6.2'), that is sixth query issued, second result of that query, as Dan Petrovic showed.

Two rival companies, two architectures, the same engineering decision: a citation is not an address, it is a coordinate in an ordered cache. What was put in store at retrieval time determines what can be cited at writing time. The rest, your page quality, its freshness, only enters the equation if it survived the previous step.

6

What gets opened gets cited

It is rare, but it changes everything. When ChatGPT actually opens a page, it cites it three times out of four.

A tenfold gap between the two measured share
A page opened by the model ends up cited74 %
A page merely present in the grounding URLs ends up cited7 %

But it is rare. Out of eighty pages pulled, only one is opened. And that opening is reserved for thinking mode on a paid account: of the 759 opens in our corpus, 757 sit there, and only two free-tier conversations open a single page.

What ChatGPT opens to reach the full page is very specific: regulatory material, official material, primary sources. Pages deep in the tree, almost never home pages.

The most opened domains in our English-language corpus: openai.com, anthropic.com, learn.microsoft.com and digital-strategy.ec.europa.eu, alongside newsrooms and pricing pages.

Evidence

Seven hundred and fifty-nine opened pages were recorded across the 1,249 answers, that is 1.2 % of the URLs in the side panel. They concentrate in 176 answers, 174 of them in thinking mode on paid accounts, that is 31 % of conversations in that regime.

The two citation rates above are not computed on those 759 pages, but on the 440 exposed as result objects: 326 cited out of 440 opened, against 4,303 out of 57,853 merely retrieved. The remaining opens are only recoverable through the references carried by citations, so they are cited by construction and would inflate the first rate without measuring anything.

A note for anyone wanting to reproduce this: since the result_source field disappeared, opened pages are no longer exposed as result objects. They can still be found through the references carried by citations, which stays reliable for telling whether the model opened pages and which ones, but underestimates the total. A page opened then discarded leaves no trace at all.

Some URLs are cited from memory

In instant mode, 85 % of citations point to a URL that appears nowhere in the retrieval net, against 31 % in thinking. Some of it comes from widgets — maps, product cards, entity cards — which carry their own links outside the search channel. The rest looks like brand homepages written from the model's own memory. We cannot yet tell the two apart.

The arXiv case, and the warning that comes with it

arXiv came up more than two thousand six hundred times in our corpus. It is cited ten times. It is never opened. One might think this is because it arrives without a snippet, and it almost always does, but the counter-test is decisive: on the one hundred and seventeen occasions where arXiv did arrive with a snippet, it was cited exactly zero times. It is not ignored because it is bare, it is simply ignored.

The same pattern holds for Reddit, massively pulled and barely displayed. ChatGPT reads preprints and Reddit threads to form a view, and only exposes ordinary web pages to the reader.

Be wary of visibility rankings

When a study lists arXiv or Reddit among the most visible sites in ChatGPT, it means the tool behind it does not tell a pulled page from a cited one. Those domains are massively pulled and practically never shown to the user. This is exactly the distinction the pyramid in the previous section measures, and it is what separates a visibility measurement from a potential-traffic one.

7

Five verticals, five separate systems

What holds for web search holds for none of the other surfaces. Each has its own supply chain, and four out of five have no connection to the engines at all.

Web search

A title and two hundred characters

The foundation, covered in the previous sections of this study. In-house index in free mode, scraped Google in thinking mode. It is the only surface where work on your page counts directly.

News

Your article, rewritten by a machine

Two regimes under one label. Through the in-house index, ChatGPT receives a summary of about eleven hundred characters, six to eight times longer than the scraped snippets served by the other route. That summary is capped at two hundred words, with a hundred-word step on a few hosts: the same cap measured on the API side for licensed publishers.

That long summary is not a licensing perk: hundreds of outlets with no agreement get the same treatment. And it is genuinely rewritten, not extracted: none of them starts with the article headline, none can be found verbatim anywhere.

Local and maps

The id format gives the source away

Here the copy does not come from your site at all, but from feeds: Yelp, TripAdvisor, Google Maps. Each provider is recognisable by its id format.

The most telling fact: nearly one in five Yelp-labelled records carries a link copied from the business's Google Business profile, while Yelp itself shows the bare address. OpenAI does the stitching. A provider label does not tell you where each field came from.

Shopping

No connection to web search

The product carousel does not share a single address with the engines. Ranking well in ChatGPT's links does nothing to get you into the product cards: these are two different jobs.

Amazon is entirely absent from the offers, including on questions that name it. And the product card is an object cached for weeks, identical for everyone: being in a product's first lookup is worth more than ranking well afterwards.

Images

Served without ever touching your server

Your images are re-hosted at OpenAI. They appear in answers without a single request reaching your infrastructure, so without any trace in your logs.

A counter-intuitive detail: in thinking mode, thumbnails transit through Bing's image delivery network. It is the only Microsoft fingerprint in our entire corpus.

The takeaway from this section

Only one of these five systems is wired to classic search. The other four are fed by feeds, catalogues and dedicated indexes, where the SEO techniques specific to each vertical are what apply: image SEO, local SEO and Google Business profiles, shopping feeds. A ChatGPT visibility strategy that only handles web search leaves four doors shut.

A question worth asking

On news, the full chain deserves to be looked at squarely. OpenAI produces a summary of your article, stores it, serves it to its users for days, and the article itself is cited only about one time in eight. Value is captured at ingestion. It is not returned at display. We lay out the facts, everyone can draw their own conclusions.

8

What we still do not know

Three open questions, each with a protocol either running or ready to launch.

  1. By what route does a new page enter OpenAI's index? We know how ChatGPT fetches a page it already knows. We do not know what gets a page in there. On a domain monitored for sixteen days, not one page entered. Four triggers are being tested in parallel: orphan page, page read once, page linked from a known page, page listed in the sitemap. The answer may well be that none of the four is enough. One result is already in, but it is about crawling, not about the index. GPTBot downloads our sitemap roughly once a day and followed none of the URLs listed only there: zero pages out of eight. That same morning, it crawled all eight pages reachable through a link, including links served to verified OpenAI robots only. Bingbot, ClaudeBot and GoogleOther did the exact opposite: they came through the sitemap and nothing else. At OpenAI, discovery runs on links. Crawling is still not indexing, though.
  2. Does asking for a long answer make it search more? In thinking mode, the number of citations clearly tracks answer length. In instant mode, the relationship is nil. We do not know whether a length instruction causes a bigger search budget, or whether hard questions simply call for both at once.
  3. Why so many snippet-less citations in instant mode? One citation in two in instant mode arrives with no snippet, against under 1 % in thinking. We have two leads. The first, dominant in our measurements: the URL comes from no search at all, it is written from the model's parametric memory, and 95 % of those citations appear nowhere in the retrieval net. The second, marginal: some channels carry no snippet at all, forum, scientific and video, but they account for only 5 % of cases. That leaves 5 % of URLs genuinely retrieved and served with an empty snippet, which neither lead explains. We cannot settle it.
An open challenge: the bare URLs that get cited

Here is the question we cannot crack, and on which we would like independent measurements.

Some search results arrive with a title, a URL, and nothing else. Not even a truncated snippet: an empty one. Set aside straight away YouTube, Reddit and arXiv, which almost never carry one. What is left are ordinary pages on ordinary domains: ESPN, Yahoo Sports, the Guardian, business directories.

You would expect the model to ignore them, having nothing to judge them on. The opposite happens. In instant mode, a snippet-less result is cited 14.9 % of the time against 8.2 % for one that carries a snippet: cited nearly twice as often while the model knows twice as little.

Our leads are not enough. On local questions, the absence of snippets is total per conversation and coincides with a map being displayed, as if the link list were mere scenery and the map carried the answer: that would explain the citation, not the missing snippet. On news questions the mechanism is different, the absence sticks to the URL and follows the same page from one conversation to the next.

If you can reproduce the phenomenon, or have an explanation, write to us. The protocol is short: capture the server stream in instant mode, isolate results of type search, set aside YouTube, Reddit and arXiv, then compare the citation rate with and without a snippet. Thinking mode does not compare the same way, it pulls so many URLs that the citation rate loses its meaning there.

9

Odds and ends we also found

Observations that did not deserve their own chapter, but did deserve to be written down.

OpenAI writes Bright Data's name in its own code. Product carousels carry an instrumentation block naming the server-side experiment allocated to the shopping engine. On three conversations from 22 July, that identifier contains the vendor name in the clear. "analytics_meta": { "ab_test.search_engine_all.layer_name": "sonic_search_engine_all", "ab_test.search_engine_all.group_name": "", "ab_test.search_engine_all.allocated_experiment": "shopping_brightdata_new_feed_serving_week_2026_06_22" } This name appears nowhere in the configuration sent to the browser: the allocation is purely server-side, and it only leaks through that instrumentation block. Set against the result_source field, one of whose four values is precisely bright, the simplest reading is that this engine is Bright Data.
Routing a given prompt is not deterministic. The same question, two days later, on the same account, can switch engine entirely. A third of replayed prompts flip outright. No conclusion holds on a single capture.
Freshness words have no effect. Writing today, right now or this weekend into your question changes nothing measurable. What triggers a live search (via bright) is the vertical and the actual volatility of the fact requested.
Text hidden with CSS is read anyway. On an on-demand read, everything hidden by a style rule is passed to the model. We document this as an observed capability, not as a tactic to use.
Your structured data is not transmitted. On an on-demand read, JSON-LD is stripped from the document before it reaches the model. A caveat though: this says nothing about what the index does with it, and that is a different step.
Your pages reach the model as Markdown. On an on-demand read, the HTML is converted to Markdown before reaching the model. An important nuance: at that stage the conversion cleans nothing, the menu and the footer come through with everything else. It is at indexing time, the step that builds the snippet, that boilerplate is stripped. And what is kept, both in the index and in the read cache, is not HTML either: pages are stored there as Markdown, already converted.
Past four megabytes, a page is not read at all. No partial read, no truncation: the page is rejected outright. In practice almost no page reaches that size.
Your best citations arrive with no tag. ChatGPT appends utm_source=chatgpt.com to 95 % of the links it displays, and in instant mode to every single one of them. One family systematically escapes it: the pages the model opened and read itself never carry that parameter. In other words, the citations worth the most, the ones coming from a page actually read, are precisely those your analytics tool will fail to attribute to ChatGPT.
An answer cache short-circuits search. The server configuration describes an instant-answer cache, valid for twenty-four hours, queried in under half a second, and triggered by semantic similarity. If your question is close enough to one already asked, no search happens at all. Details here.
10

Method: how we know

We measured what ChatGPT did with our pages.

1,249
answers with retrieval
88,000
search results
26,900
distinct pages pulled
6,400
distinct domains
The corpus

One thousand two hundred and forty-nine ChatGPT answers captured in real conditions. Six hundred and eighty-two in instant mode and five hundred and sixty-seven in thinking mode. A little under five hundred on free accounts, a little over seven hundred and fifty on paid accounts.

The counting unit throughout is the answer-and-page pair. The same page seen in two answers counts twice, the same page seen twice inside one answer counts once. Addresses are normalised before counting, tracking parameters stripped.

Why twelve hundred answers beat a million

These are neither API calls nor interface monitoring. They are real conversations on real accounts: free, paid, several distinct paid accounts, in several countries, and even logged out.

Most market tools watch the logged-out interface. They therefore observe a single regime, and one of the least representative ones at that. It is precisely the spread between our accounts that isolated the role of each engine: without a free-versus-paid comparison at identical prompt, there is no way to see that routing depends on the model-and-effort pair rather than on the subscription.

The six instruments

The first four are ours. The last two come from Jérôme Salomon (Oncrawl), and look at the same machine from the other end: the server that receives the robots, and the API that answers off the same engine as the application.

01

A lab of tagged pages hosted on one of our own sites. Every page served carries a random fingerprint, written at the same instant into the page and into our log. So we know which precise visit produced the copy the model spits back, and when.

02

A Chrome capture extension we build and distribute freely. It records the raw server stream of each answer, before anything is rendered on screen.

think.resoneo.com/scrap-chatgpt-plugin

03

A fleet of ChatGPT accounts: free, paid, several distinct paid accounts, plus logged-out mode. All of it behind a VPN, to switch address and country at will.

04

A feature-flag tracker that archives, every day, the configuration embedded in the public HTML of chatgpt.com. It dates server-side switches to the day.

think.resoneo.com/chatgpt-experiments

05 · Oncrawl

Server log analysis, over three months and 325,000 URLs. Every GPTBot, OAI-SearchBot and ChatGPT-User visit is dated there, at the scale of a whole site. That is what makes it possible to say which robot came when, to measure each one's pace, and to date a cache retention on a log line rather than on an inference.

06 · Oncrawl

The search API turned into an instrument. On POST /v1/responses with the web_search tool, the include: ["web_search_call.results"] parameter surfaces the internal metadata the application never displays: Crawled, the age of the served copy URL by URL, and wordlim, the word budget granted to the domain. And external_web_access: false queries the store without going out to the web. 7,782 results analysed.

detail in the collaboration section

The protocol, then cross-validation

The conversations are not taken at random. The same prompts, word for word, were replayed on every account type. That is the only way to separate what comes from the subscription from what comes from the answer mode. And the prompt sets are split by theme, to cover every surface rather than just one: general search, news, shopping, local and maps, images, sport, culture, admin and health.

Only then, to find out which engine sits behind each pipeline, we replayed the queries ChatGPT writes for itself, the famous fan-outs, on Google and on Bing, then looked for the URLs it had served in both engines' results.

A small irony we own: to find out whether OpenAI scrapes Google, we scraped Google. Same weapons ^^

Our March study described the tool and its commands. This one describes what flows through it. think.resoneo.com/chatgpt/5.3-5.4

11

Credits, sources and tools

This work is a continuation. It would not be possible without several people who published before us, and who published their method alongside their results.

Suganthan
Mohanadasan
Put the field naming the engines into the public domain, by reading network traffic rather than answers, and established its four values. That is the starting point of this work, including the one place where we contradict him: our measurements do not support reading labrador as a licensed-content catalogue.
Mark Williams-Cook
Metehan Yeşilyurt
Put us on this track. The first documented that ChatGPT quietly runs real Google searches and made those queries readable, which our cross-validation confirms by another route. The second continues in-depth work on ChatGPT's internal ranking.
Chris Green Provided the only large-scale public dataset on engine distribution, close to ten thousand search runs. Our own distributions read against his.
David Konitzny First documented ChatGPT's local infrastructure and spotted the anonymised provider identifiers. We start from his observation to show that those identifiers share the schema of the named feeds, that one of the ones he lists never materialised in our captures, and that a fourth provider, openly Google, fills the categories Yelp covers poorly.
Jérôme Salomon Guided us through the whole cache section, and uncovered the parameter that lets you query OpenAI's cache with no web access, concrete proof that the discovery index and the read cache exist. We also owe him three measurements we lacked: OAI-SearchBot feeds the store, the noindex directive is ignored inside it, and being read does not guarantee admission. Plus the retention beyond ninety days, which we did not reproduce and report as such. And it was the hunch of instrumenting the web_search API that opened the whole internal-metadata investigation. The detail is in the collaboration section.
Dan Petrovic Exposed Gemini's internal anchoring format and its query-by-result decoding. The parallel with ChatGPT's anchors structures section 5.
What the new version of the extension brings new

The extension that produced this corpus has just been updated. It no longer merely records the raw stream: answer by answer, it rebuilds the path of a URL from the bottom of the side panel up to the citation pill. The figures in this study can therefore be recomputed on your own conversations.

New version of the capture extension, listed, promoted, cited funnel
The funnel as it appears in the extension, on a single conversation. Each result nature carries its own colour.

Get the capture extension

  • The listed, promoted, cited funnel, per conversation. The URLs in the side panel, then the ones pulled up to the top section, main source and supporting sources, then the lead source carried by the citation pill. In absolute values and as a share of the panel. Carousels are counted separately.
  • An aggregated pyramid on the dashboard. Cumulated over the filtered conversations, with a result-nature selector: panel, top, lead source and opened pages for each nature, plus the promotion rates inside that nature.
  • Eleven named, colour-coded result natures, in place of the raw ref_type: web search, news, scientific, forum, video, local, product, image, weather, opened page, and no ref_type.
  • Explicit labels for link type: citations at the top of the panel, other panel sources, footnote, widgets, opened and therefore fetched, and the residue.
  • A card for the pages opened by the model. It is built on the Total lines: fetch marker and on the read references, deduplicated by normalised URL, with the fetch signature kept.
  • A labrador badge. Since result_source vanished from the server stream, between 20 and 22 July, results from OpenAI's in-house index are inferred from format signatures: a snippet of at least one hundred and eighty-six characters, an all-caps snippet header. The signals used are shown in the tooltip, so the inference stays checkable.

Other sources and studies

July 2026 ChatGPT experiments tracker: what OpenAI tests before rollout July 2026 How your Google phone tracks you to "serve you better" June 2026 3,729,456 Google internal URLs, without opening a single one June 2026 What Google is really building June 2026 Inside Pinterest's algorithm June 2026 How Chrome classifies websites internally May 2026 Tomorrow's AI phone, seen from inside a Google APK May 2026 Ranking of the top Google Preferred Sources Apr 2026 Inside Brave Search: the invisible infrastructure of genAI Apr 2026 How ChatGPT Search works? Full reverse engineering More stuffs...
RESONEO