27,000 pages put under the microscope. The engines behind the answers, the two hundred characters that stand for you, and the caches that freeze them.
When ChatGPT answers using the web, it does not go and read your page. It queries one or more engines, takes away a title, an address and roughly two hundred characters, and works from that. The page itself is opened in barely more than one case out of 80, and in thinking mode only. The rest of the time, what stands for you in front of the model fits in three lines of text produced elsewhere, at another moment, and sometimes served to someone else before you.
This study takes that chain apart layer by layer, across more than twelve hundred real answers captured in July 2026. Who brings the results back, what the model sees of them exactly, how long that copy stays valid, and what clears the final step to become a visible link in the answer.
In the code, every result used to carry the engine that brought it back, through the result_source field. That field has vanished from the stream. We are working on a window that has just closed.
deprecated.The result_source field was carried by every result object in the server stream, with four possible values: labrador, bright, oxylabs, serp. Across the 1,249 answers in the corpus, 760 carry it and 489 do not. Last labelled stream on 21 July, first fully bare day on 22 July.
Two of the four engines, bright and oxylabs, serve scraped Google. We established it by replaying the same queries on both engines, then looking for the URLs used by ChatGPT in their results. One URL in three shows up on Google's first page. Only one in twenty on Bing. And when it is in Google, nine titles out of ten match character for character, ellipsis included, in the same spot.
What makes the gap conclusive is that Google and Bing barely overlap with each other on these queries. Being present in one and not in the other therefore becomes a strong signal rather than a ranking coincidence. As for labrador, it belongs to neither: it is an in-house index, not a scrape.
The same move shows up on the shopping side. Google tokens have left the product carousels, but the Google product id survives in the clear inside the request fired when a user clicks a card, this time alongside an encrypted block.
And the dependency is not residual, it is functional. The offers displayed in the carousels do come from OpenAI's internal catalogue, fed by merchant feeds. But that catalogue cannot do everything: Google remains the oracle on best price, on part of the product metadata and on reviews. In other words, OpenAI has made its dependency unreadable in the visible stream, while keeping it operational in the layer that only fires on click.
The removal falls two days after Google's complaint against SerpApi was dismissed. We flag it as a calendar match, not as a cause-and-effect mechanism. Nothing in our measurements links the two, but you have to admit it is unsettling ;)
The free tier does not reason. Across our 493 free-tier answers, not a single one went into thinking mode. And instant settles for a single fan-out in more than eight cases out of ten, and never opens a page to read it.
Median values per answer. Thinking mode reaches nine times wider, and sometimes up to several hundred pages on a single question.
In instant mode, on a free account, four families of questions and only two behaviours.
| What you ask for | labrador | No snippet | |
|---|---|---|---|
| Stable information | 100 % | 0 % | 1 % |
| A business, a place | 100 % | 0 % | 25 % |
| A product to buy | 100 % | 0 % | 1 % |
| Live information | 50 % | 50 % | 8 % |
The last row is the questions whose answer changes within the day: who won the game last night, the Masters leaderboard live, today's Fed decision. They are the only ones that open the door to scraped Google. Everything else, and that is most of what people ask ChatGPT, never leaves the in-house index.
In other words, understanding labrador is not a niche topic: it supplies almost every result, for the population that actually uses ChatGPT.
Search results only, empty snippets excluded from the share, across 5,245 snippets and 424 free answers in instant mode. Stable information: 99.9 % labrador across 2,136 snippets. Business and place: 100 % across 310. Product: 97.6 % across 1,392. News: 52.6 % labrador and 46.1 % Google across 1,407. The split is read off the length of the served snippet, which separates the two engines unambiguously: past one hundred and ninety-five characters the source is labrador in 99.9 % of cases, measured over the period when the stream still named its supplier.
The missing-snippet column deserves to be read alongside the first: on businesses and places, a quarter of results arrive with no snippet at all, just a title and an address. The best-served class in the table is also the one where the model receives the least material.
The table on the left is for instant mode. Ask the same questions again in thinking and the ratio turns over, across all four families at once.
| The same question, in thinking | labrador | |
|---|---|---|
| Stable information | 1 % | 99 % |
| A business, a place | 3 % | 95 % |
| A product to buy | 1 % | 99 % |
| Live information | 0 % | 96 % |
In instant, the family of the question decides everything. In thinking, it has no effect at all: the four values close up where instant spreads them from 50 to 100 %. Businesses and places even change architecture between the two regimes. In instant, the answer is carried by a map, and the result list is just a reminder of links, a quarter of which have no snippet at all. In thinking, the map disappears entirely, the list goes back to being a real retrieval net, and the share of snippet-less results drops from 25 % to 4 %.
Two readings of labrador circulate, and our measurements rule out both. The first makes it Bing served under another name, on the grounds of the OpenAI-Microsoft partnership. One measurement is enough to overturn it: Bing cuts the titles it displays at seventy-five characters, never one more, and a quarter of labrador titles run past that ceiling. The longest one we have seen runs to nine hundred and seventy-seven.
The maximum length of title tags displayed in Bing's result pages is seventy-five characters, never one more. That ceiling is what ruled out the hypothesis that labrador is rebadged Bing: a quarter of labrador titles cross it, which is physically impossible for a Bing scrape.
A fifth engine, named bing, has been reported by other researchers. We have never observed it, on any account, in any regime, on any date. Our hypothesis is that it may be exposed only to OpenAI Teams accounts, because of the Microsoft relationship.
Jérôme Salomon has passed us June captures where that pipeline shows up in bulk on a Team account, while the matching free account sees none of it on the same day (collaboration section).
Two other measurements point the same way. Snippet lengths are incompatible: Bing cuts wide and irregular, labrador stops dead at two hundred and two characters, across every one of its results. And the labrador snippet is no engine's snippet: on labrador URLs that also rank on Google and on Bing, text reuse is nil on both sides. The same measurement applied to bright finds Google's text four times out of ten.
The second reading makes labrador the preserve of publishers under agreement with OpenAI. Our measurements do not support that one either. At equal engine, having a licensing deal or not changes nothing about the treatment: hundreds of outlets with no agreement get exactly the same format, the same length, the same freshness as the partners.
Labrador is better described as a repository of immediate access.
What defines it is not the contract, it is that OpenAI queries it without paying a third party and without waiting for a scrape.
All engine shares in this section are computed on results of type search only, across the 373 answers before the cut-off, and read straight from the result_source field, with no classifier.
This filter is essential. Counted across all types, oxylabs looks like 12.7 % of instant mode. On search results alone it drops to zero: oxylabs is a news pipeline and nothing else, all of its results are of type news. Symmetrically, scientific results and forums are served 100 % by labrador, so including them mechanically inflates its share.
Exact values, search results only. Free and instant: labrador 63.6 %, bright 34.7 %, serp 1.6 %, across 1,793 results and 158 answers. Paid and thinking: bright 74.6 %, labrador 23.6 %, serp 1.9 %, across 16,407 results and 176 answers.
Both averages hold, but they depend entirely on what is being asked. The batch that produced them was half made of news questions, the one family that goes to Google. On a batch of stable questions, labrador's share on the free tier rises to 100 %. An engine average therefore only means something alongside the mix of questions that produced it: it is the table by family, above, that describes how the system behaves.
In instant mode, everything fits in a title and two hundred characters. And those two hundred characters are not the ones you think.
Once its fan-outs are fired, ChatGPT receives a list of results exactly like the one you get when typing a query into a search engine: for each page, a title, a snippet and a URL. Nothing more. It is on that grounding net, and on it alone, that the model relies to write its answer and decide who it will cite.
In other words, those three fields are your only representation in front of the model in the vast majority of cases. Hence the importance of knowing exactly how they are built.
Depending on the engine that brings it back, the same page arrives with its full title or with a title already truncated by a results page. Through labrador, the tag passes whole: a quarter of titles run past seventy-five characters and almost none end in an ellipsis. Through scraped Google, the title is cut to an engine's display length, and nearly one in four carries the mark of that cut.
Here are the three formats actually passed to the model, from shortest to longest. The same structure every time, a URL, a title and a snippet, but budgets that vary sevenfold.
The snippet stored by labrador's in-house index is cut off just past two hundred characters, and it is taken from the start of the page's rendered body, not from your metadata. Your meta description is ignored. It stays useful for bright results, the Google scrape, which picks it up one time in three.
The cut almost never falls in the middle of the H1: across nearly four hundred snippets containing it, only four truncate it. Your grounding budget is therefore perfectly predictable. It is made of three blocks that always follow the same order: whatever sits before the H1, the H1 itself, and whatever room is left afterwards for your actual content.
What eats the start of the snippet is not what people assume. Breadcrumbs and bylines are marginal. The two real losses are the section kicker and the alt text of the first image, the one sitting right after the title, which on its own can eat fifty characters.
| What comes before the H1 | Frequency | Characters lost |
|---|---|---|
| Kicker, section, category | 29 % | 18 |
| Publication or update date | 11 % | 25 |
| Alt text of the first image, right under the title | 9 % | 50 |
| Site or brand name | 7 % | 44 |
| Breadcrumb | 7 % | 23 |
| Author byline | 3 % | 45 |
The first lever is not there, though. Across the pages we examined, one in seven simply has no H1 markup at all. In that case the anchoring becomes unpredictable: the snippet starts on some subheading picked by the template, and you no longer control anything.
Method: 534 pages actually cited by ChatGPT were fetched then compared with the snippet the index had stored for them. Of the 463 that have an H1, 387 snippets contain it, that is 83.6 %. Among those 387, the H1 is the very first character in 297 cases, that is 77 %. The prefix, when there is one, is 20 characters at the median.
The snippet is not chosen based on the question asked: seven regional variants of the same press release, queried with the same request, produce seven different anchor points. And it is not built at search time: the vast majority of pages seen under several different queries return a strictly identical snippet. It is frozen at crawl time, and it ages.
What to do with this. Your grounding budget in instant mode is your full title plus roughly one hundred and fifty useful characters starting at your H1. Mark up an H1, clear the runway between it and the first paragraph, and keep an eye on the alt text of the first image, the one sitting right after the title.
OpenAI keeps your pages in two distinct stores, written by two different robots. The copy you have in there has an age. You do not choose it, and you do not see it in your logs.
Everything else in this section reads through this one distinction. The index is what finds you, the read cache is what reads you. Not the same content, not the same robots, not the same delays.
| The index | The read cache | |
|---|---|---|
| What it holds | a short snippet, frozen at crawl time | the whole page, already converted to Markdown |
| You reach it through | a query, some keywords | an address, the exact URL |
| What it is for | finding your page | reading it |
| Who writes it | OAI-SearchBot | ChatGPT-User, and OAI-SearchBot too |
| What updates it | a SearchBot visit, and nothing else | a user request, once past thirty minutes |
| Lifetime | no expiry observed; 13 % of snippets run more than a month behind the crawl | no expiry observed; several weeks, if not months |
| How to measure it | the API Crawled field, which leaves no trace in your logs | ChatGPT-User showing up in your server logs |
The two do not hold the same pages. Jérôme found that not every URL in the index turns up in the read cache. Being findable and being readable are not the same thing.
Below that window, the page is not reopened at all. One user asks for your page at 2 pm, another asks for it at 2:29 pm: the 2 pm copy is what the second one gets, and ChatGPT does not even bother going back to check whether the page changed.
With no new request, the copy is not discarded. We observed no expiry.
The mechanism is a cache that serves first and checks afterwards. Once the copy is past the half hour, ChatGPT still shows you the old version, then goes off to refresh in the background. The fresh version will only be served on the next request. In other words, when you see ChatGPT-User in your logs, it is not a read for the current user: it is a refresh of OpenAI's cache, which the next visitor will benefit from.
On our test site, a random fingerprint is written at the same instant into the served page and into our server logs. When the model spits it back, we know exactly which visit produced the copy it just read.
Here is the sequence that established the cache is shared between users. A fresh page, never read by anyone, and two paid accounts with no connection between them.
The copy has just been made.
It hands back, word for word, account A's fingerprint.
Account B never touched our server. It read account A's copy.
Double proof: the fingerprint is identical, and the hit counter stays at zero. Either one alone would prove nothing.
The copy created by one account is served back to another account, in another country, on another plan, without a single hit reaching the origin server. The cache key is the page address, not the user.
Another point worth knowing: our test pages sent a Cache-Control: no-store header, which explicitly forbids any caching. They get cached anyway.
ChatGPT-User builds the read cache, the whole-page one: it is the robot that fires when the model or the user explicitly asks for a page. OAI-SearchBot, the search robot, builds the index, the snippet one, and it also feeds the read cache, though not systematically. Pages it visited alone, with no on-demand read ever fired, are indeed in there.
The reverse is not true, and that is the counter-intuitive part. ChatGPT-User never touches the index. Across pages dated both by OpenAI's index and by server logs, the served snippet tracks the SearchBot visit within a few days; yet ChatGPT-User is the more recent of the two robots in most cases. So it does come back to read your pages, it does refresh its whole-page copy, and none of that changes the snippet search displays.
Almost every index entry is more recent than GPTBot's last visit, so refreshed after it came and not by it, and we had verified the same thing on the read cache. The practical consequence is worth stating: being crawled and being available are two different things. The detail of that measurement sits in the collaboration section.
Beyond the two main stores, ChatGPT also keeps objects it built itself. They are not read like pages, but they age the same way.
| What is cached | Observed lifetime | Consequence |
|---|---|---|
| The summary of a news article | stable across captures | Generated once, served repeatedly afterwards |
| The product or entity card (ChatGPT sidebar) | weeks or months | Identical from one account to the next, across countries too |
On the product card, the demonstration is striking: the same product clicked from a French paid account and from an American free account returns exactly the same text, generated once.
We bracketed the freshness window by dichotomy on two different pages and two different days, by reading the served fingerprint while checking in parallel for the presence or absence of a server hit. No refresh at 3, 16 and 25 minutes. Refresh at 31 and 33 minutes, then at 1 hour, 9 hours and 10 days. The threshold therefore sits between 25 and 31 minutes, and 30 is a plausible round value rather than a value we read anywhere.
Retention is a lower bound: a copy from 11 July was still being served on the 23rd. We never observed an expiry, which does not prove there is none. On Jérôme Salomon's testing ground, pages were still being served from this cache more than ninety days after they were fetched, with a server log to back it.
Two investigations run in parallel on the same questions, one at Oncrawl, one here. Here is what came out of putting them side by side.
Throughout these investigations we worked in dialogue with Jérôme Salomon, technical SEO expert at Oncrawl. Each side ran its own measurements, with its own tools and its own testing ground; at every exchange, one set of findings corrected or extended the other. Here is what that collaboration established, between cross-confirmations and complementary finds.
The previous section owes almost everything to this part of the dialogue. Two independent protocols answer each other here: server logs and the API on one side, watermarked pages and the account fleet on the other. Five facts, established by Jérôme.
external_web_access: false to the API asks the model to answer without going out to the web. Whatever it hands back then necessarily comes from its store. That is concrete proof that the discovery index and the read cache exist, and the starting point of the whole previous section. meta robots: noindex and visited by OpenAI's robots still end up in the cache. Forbidding indexing does not keep you out of the store. Jérôme's hunch was right: the API exposes the same machinery as the application, down to its internal markers. The engine itself cannot be picked from the API, but the lead pointed at the right place: digging into it turned up the field nobody looks at.
The lever fits on one line. On POST /v1/responses, with the web_search tool, you ask for include: ["web_search_call.results"]. Every result then arrives prefixed with internal metadata the application stream never exposes.
Three pieces of information, across the 7,782 results we analysed through the API. Crawled gives the age of the copy being served, URL by URL, with no access to the site's logs whatsoever: an audit oracle that exists nowhere else. wordlim gives the word budget allocated to the domain. And the citation marker names the source channel: search, news, science, forum, opened page. None of it is documented, and nothing stops OpenAI from removing it tomorrow, the way the field naming the pipelines was removed on 21 July.
| Budget per source | Who it applies to | Reading |
|---|---|---|
| 25 words | Guardian, ESPN, Washington Post, Bloomberg, Corriere, Spiegel… | Cause undetermined |
| 100 words | 11 domains identified, all of them known licensees: Le Monde, WSJ, Politico, Bild, Condé Nast, Hearst… | Looks like a contract clause |
| 200 words | The default, all the rest of the web | The normal regime |
The value is a constant per domain-and-channel pair, with no residual variance at all once the right key is found, and a domain's PDFs inherit the domain's budget. The 100 class explains itself. The 25 class stays a puzzle: it mixes publishers with an agreement and publishers without one, and the robots.txt explanation is refuted by counter-examples in both directions.
That word cap is not an API invention. It sat in the ChatGPT system prompt we published in March: the [wordlim N] instruction, with two hundred words per source by default, is present in the GPT-5.3 version. The API now confirms it server-side, source by source. The loop closes between what the prompt ordered yesterday and what the stream measures today.
Excerpt from the system prompt, published in our March study.
The age oracle, for its part, holds two surprises for the general reader. On the live web it does date a real read, decoupled from the publication date. But on corpora ingested in bulk, it dates the content of the copy, and that copy was never refreshed.
A Wikipedia page we analysed, edited several times a day, is served with a "Crawled: 2 months ago" marker. Across 308 Wikipedia results, 209 are dated between eight and ninety days: the encyclopedia models cite most is served from a periodic snapshot, not from the site. Those ages fit the encyclopedia's own rhythm, whose full exports are produced once a month: the copy follows the dump schedule, not the edit schedule.
On Reddit, the collection date is the date the thread was created, zero days apart at the median, going back as far as seven years. A thread opened in 2019 is served as it stood at creation, or a few weeks later: replies added after that are not in the copy.
This is the ground where our two observation posts differ the most, and overlap the best. Three points.
site:reuters.com queries, all of them empty. So there is a direct ingestion channel for press content, one that does not depend on the crawl. Jérôme supplies the mechanism: partners have an API to push their articles to OpenAI, and those articles are available without ever having needed a first crawl. They do get crawled afterwards, by ChatGPT-User and by SearchBot, just less than the articles that do not go through that channel. The concrete consequence for publishers: the best integrated are also the least crawled, and no ChatGPT in your logs does not mean no ChatGPT in the answers. news flag; picked up by the crawl, it arrives with the search flag. A 30 July capture shows both sides in the same answer, for the same publisher: Le Monde appears seven times as news with a 180-word summary, and once as search with a 32-word snippet. The flag does not describe the topic, it describes the door. bing pipeline in bulk, on 93.7 % of the sources brought back. That same month, on our accounts, zero. We ruled out two explanations: it is neither the server region that handles the answer, nor free versus paid. What remains is the Team plan: the most likely hypothesis. We suspect an arrangement between OpenAI and Microsoft opening Bing access to ChatGPT Team users, without being in a position to verify it. Everything ChatGPT pulls ends up listed somewhere. But between being listed and being cited there are three steps, and almost nothing climbs them.
Nothing in the data marks a result as retained. The notion of a citation only exists at display time, which is why published figures on the subject diverge so much from one study to the next: each counts at a different layer without saying so. We rebuilt the layers from the raw stream, then checked how they map onto the interface across two real sessions where the client names each display zone itself.
It is tempting to write that most pulled pages are never seen. That is false, and we believed it for a while. Every page brought back by retrieval is indeed listed in the panel that opens when you click "Sources" under the answer. Across more than sixty thousand answer-and-page pairs, not one is absent from every display surface.
What is true is subtler: nine out of ten stay at the bottom of that panel, in the "More" section, behind one click, never promoted to the top and never picked up in the text. Technically visible, practically invisible.
A citation anchor in the text can carry several sources. The pill shows the name of the first one and flags the others with a "+1", reachable by opening the tooltip. All of them are attached to the citation, but only one gives it its name.
That is the whole gap between the two steps below: roughly 1/3 of URLs are attached to a citation without being its lead source. They exist, they are clickable, but they are not what the reader sees first.
The first column follows the Sources panel, step by step. The second column says what the model actually did with the pages. The filters cross the mode, the nature of the result and the engine; combinations with too little data are flagged.
A median answer lists around twenty pages at the bottom of the sources side panel, promotes five to the top, cites three in its text and opens none. The median for opened pages is zero: nine answers out of ten never open anything.
Nine pages out of ten never leave the bottom of the panel. They were selected, passed to the model, they potentially contributed to the answer, but ChatGPT gives them almost no prominence.
ChatGPT encodes its citations as turn0search11: turn zero, search family, eleventh result. Gemini encodes its own as PerQueryResult(index='6.2'), that is sixth query issued, second result of that query, as Dan Petrovic showed.
Two rival companies, two architectures, the same engineering decision: a citation is not an address, it is a coordinate in an ordered cache. What was put in store at retrieval time determines what can be cited at writing time. The rest, your page quality, its freshness, only enters the equation if it survived the previous step.
It is rare, but it changes everything. When ChatGPT actually opens a page, it cites it three times out of four.
But it is rare. Out of eighty pages pulled, only one is opened. And that opening is reserved for thinking mode on a paid account: of the 759 opens in our corpus, 757 sit there, and only two free-tier conversations open a single page.
What ChatGPT opens to reach the full page is very specific: regulatory material, official material, primary sources. Pages deep in the tree, almost never home pages.
The most opened domains in our English-language corpus: openai.com, anthropic.com, learn.microsoft.com and digital-strategy.ec.europa.eu, alongside newsrooms and pricing pages.
Seven hundred and fifty-nine opened pages were recorded across the 1,249 answers, that is 1.2 % of the URLs in the side panel. They concentrate in 176 answers, 174 of them in thinking mode on paid accounts, that is 31 % of conversations in that regime.
The two citation rates above are not computed on those 759 pages, but on the 440 exposed as result objects: 326 cited out of 440 opened, against 4,303 out of 57,853 merely retrieved. The remaining opens are only recoverable through the references carried by citations, so they are cited by construction and would inflate the first rate without measuring anything.
A note for anyone wanting to reproduce this: since the result_source field disappeared, opened pages are no longer exposed as result objects. They can still be found through the references carried by citations, which stays reliable for telling whether the model opened pages and which ones, but underestimates the total. A page opened then discarded leaves no trace at all.
In instant mode, 85 % of citations point to a URL that appears nowhere in the retrieval net, against 31 % in thinking. Some of it comes from widgets — maps, product cards, entity cards — which carry their own links outside the search channel. The rest looks like brand homepages written from the model's own memory. We cannot yet tell the two apart.
arXiv came up more than two thousand six hundred times in our corpus. It is cited ten times. It is never opened. One might think this is because it arrives without a snippet, and it almost always does, but the counter-test is decisive: on the one hundred and seventeen occasions where arXiv did arrive with a snippet, it was cited exactly zero times. It is not ignored because it is bare, it is simply ignored.
The same pattern holds for Reddit, massively pulled and barely displayed. ChatGPT reads preprints and Reddit threads to form a view, and only exposes ordinary web pages to the reader.
When a study lists arXiv or Reddit among the most visible sites in ChatGPT, it means the tool behind it does not tell a pulled page from a cited one. Those domains are massively pulled and practically never shown to the user. This is exactly the distinction the pyramid in the previous section measures, and it is what separates a visibility measurement from a potential-traffic one.
What holds for web search holds for none of the other surfaces. Each has its own supply chain, and four out of five have no connection to the engines at all.
The foundation, covered in the previous sections of this study. In-house index in free mode, scraped Google in thinking mode. It is the only surface where work on your page counts directly.
Two regimes under one label. Through the in-house index, ChatGPT receives a summary of about eleven hundred characters, six to eight times longer than the scraped snippets served by the other route. That summary is capped at two hundred words, with a hundred-word step on a few hosts: the same cap measured on the API side for licensed publishers.
That long summary is not a licensing perk: hundreds of outlets with no agreement get the same treatment. And it is genuinely rewritten, not extracted: none of them starts with the article headline, none can be found verbatim anywhere.
Here the copy does not come from your site at all, but from feeds: Yelp, TripAdvisor, Google Maps. Each provider is recognisable by its id format.
The most telling fact: nearly one in five Yelp-labelled records carries a link copied from the business's Google Business profile, while Yelp itself shows the bare address. OpenAI does the stitching. A provider label does not tell you where each field came from.
The product carousel does not share a single address with the engines. Ranking well in ChatGPT's links does nothing to get you into the product cards: these are two different jobs.
Amazon is entirely absent from the offers, including on questions that name it. And the product card is an object cached for weeks, identical for everyone: being in a product's first lookup is worth more than ranking well afterwards.
Your images are re-hosted at OpenAI. They appear in answers without a single request reaching your infrastructure, so without any trace in your logs.
A counter-intuitive detail: in thinking mode, thumbnails transit through Bing's image delivery network. It is the only Microsoft fingerprint in our entire corpus.
Only one of these five systems is wired to classic search. The other four are fed by feeds, catalogues and dedicated indexes, where the SEO techniques specific to each vertical are what apply: image SEO, local SEO and Google Business profiles, shopping feeds. A ChatGPT visibility strategy that only handles web search leaves four doors shut.
On news, the full chain deserves to be looked at squarely. OpenAI produces a summary of your article, stores it, serves it to its users for days, and the article itself is cited only about one time in eight. Value is captured at ingestion. It is not returned at display. We lay out the facts, everyone can draw their own conclusions.
Three open questions, each with a protocol either running or ready to launch.
Here is the question we cannot crack, and on which we would like independent measurements.
Some search results arrive with a title, a URL, and nothing else. Not even a truncated snippet: an empty one. Set aside straight away YouTube, Reddit and arXiv, which almost never carry one. What is left are ordinary pages on ordinary domains: ESPN, Yahoo Sports, the Guardian, business directories.
You would expect the model to ignore them, having nothing to judge them on. The opposite happens. In instant mode, a snippet-less result is cited 14.9 % of the time against 8.2 % for one that carries a snippet: cited nearly twice as often while the model knows twice as little.
Our leads are not enough. On local questions, the absence of snippets is total per conversation and coincides with a map being displayed, as if the link list were mere scenery and the map carried the answer: that would explain the citation, not the missing snippet. On news questions the mechanism is different, the absence sticks to the URL and follows the same page from one conversation to the next.
If you can reproduce the phenomenon, or have an explanation, write to us. The protocol is short: capture the server stream in instant mode, isolate results of type search, set aside YouTube, Reddit and arXiv, then compare the citation rate with and without a snippet. Thinking mode does not compare the same way, it pulls so many URLs that the citation rate loses its meaning there.
Observations that did not deserve their own chapter, but did deserve to be written down.
result_source field, one of whose four values is precisely bright, the simplest reading is that this engine is Bright Data.utm_source=chatgpt.com to 95 % of the links it displays, and in instant mode to every single one of them. One family systematically escapes it: the pages the model opened and read itself never carry that parameter. In other words, the citations worth the most, the ones coming from a page actually read, are precisely those your analytics tool will fail to attribute to ChatGPT.We measured what ChatGPT did with our pages.
One thousand two hundred and forty-nine ChatGPT answers captured in real conditions. Six hundred and eighty-two in instant mode and five hundred and sixty-seven in thinking mode. A little under five hundred on free accounts, a little over seven hundred and fifty on paid accounts.
The counting unit throughout is the answer-and-page pair. The same page seen in two answers counts twice, the same page seen twice inside one answer counts once. Addresses are normalised before counting, tracking parameters stripped.
These are neither API calls nor interface monitoring. They are real conversations on real accounts: free, paid, several distinct paid accounts, in several countries, and even logged out.
Most market tools watch the logged-out interface. They therefore observe a single regime, and one of the least representative ones at that. It is precisely the spread between our accounts that isolated the role of each engine: without a free-versus-paid comparison at identical prompt, there is no way to see that routing depends on the model-and-effort pair rather than on the subscription.
The first four are ours. The last two come from Jérôme Salomon (Oncrawl), and look at the same machine from the other end: the server that receives the robots, and the API that answers off the same engine as the application.
A lab of tagged pages hosted on one of our own sites. Every page served carries a random fingerprint, written at the same instant into the page and into our log. So we know which precise visit produced the copy the model spits back, and when.
A Chrome capture extension we build and distribute freely. It records the raw server stream of each answer, before anything is rendered on screen.
A fleet of ChatGPT accounts: free, paid, several distinct paid accounts, plus logged-out mode. All of it behind a VPN, to switch address and country at will.
A feature-flag tracker that archives, every day, the configuration embedded in the public HTML of chatgpt.com. It dates server-side switches to the day.
Server log analysis, over three months and 325,000 URLs. Every GPTBot, OAI-SearchBot and ChatGPT-User visit is dated there, at the scale of a whole site. That is what makes it possible to say which robot came when, to measure each one's pace, and to date a cache retention on a log line rather than on an inference.
The search API turned into an instrument. On POST /v1/responses with the web_search tool, the include: ["web_search_call.results"] parameter surfaces the internal metadata the application never displays: Crawled, the age of the served copy URL by URL, and wordlim, the word budget granted to the domain. And external_web_access: false queries the store without going out to the web. 7,782 results analysed.
The conversations are not taken at random. The same prompts, word for word, were replayed on every account type. That is the only way to separate what comes from the subscription from what comes from the answer mode. And the prompt sets are split by theme, to cover every surface rather than just one: general search, news, shopping, local and maps, images, sport, culture, admin and health.
Only then, to find out which engine sits behind each pipeline, we replayed the queries ChatGPT writes for itself, the famous fan-outs, on Google and on Bing, then looked for the URLs it had served in both engines' results.
Our March study described the tool and its commands. This one describes what flows through it. think.resoneo.com/chatgpt/5.3-5.4
This work is a continuation. It would not be possible without several people who published before us, and who published their method alongside their results.
| Suganthan Mohanadasan |
Put the field naming the engines into the public domain, by reading network traffic rather than answers, and established its four values. That is the starting point of this work, including the one place where we contradict him: our measurements do not support reading labrador as a licensed-content catalogue. |
| Mark Williams-Cook Metehan Yeşilyurt |
Put us on this track. The first documented that ChatGPT quietly runs real Google searches and made those queries readable, which our cross-validation confirms by another route. The second continues in-depth work on ChatGPT's internal ranking. |
| Chris Green | Provided the only large-scale public dataset on engine distribution, close to ten thousand search runs. Our own distributions read against his. |
| David Konitzny | First documented ChatGPT's local infrastructure and spotted the anonymised provider identifiers. We start from his observation to show that those identifiers share the schema of the named feeds, that one of the ones he lists never materialised in our captures, and that a fourth provider, openly Google, fills the categories Yelp covers poorly. |
| Jérôme Salomon | Guided us through the whole cache section, and uncovered the parameter that lets you query OpenAI's cache with no web access, concrete proof that the discovery index and the read cache exist. We also owe him three measurements we lacked: OAI-SearchBot feeds the store, the noindex directive is ignored inside it, and being read does not guarantee admission. Plus the retention beyond ninety days, which we did not reproduce and report as such. And it was the hunch of instrumenting the web_search API that opened the whole internal-metadata investigation. The detail is in the collaboration section. |
| Dan Petrovic | Exposed Gemini's internal anchoring format and its query-by-result decoding. The parallel with ChatGPT's anchors structures section 5. |
The extension that produced this corpus has just been updated. It no longer merely records the raw stream: answer by answer, it rebuilds the path of a URL from the bottom of the side panel up to the citation pill. The figures in this study can therefore be recomputed on your own conversations.
ref_type: web search, news, scientific, forum, video, local, product, image, weather, opened page, and no ref_type.Total lines: fetch marker and on the read references, deduplicated by normalised URL, with the fetch signature kept.result_source vanished from the server stream, between 20 and 22 July, results from OpenAI's in-house index are inferred from format signatures: a snippet of at least one hundred and eighty-six characters, an all-caps snippet header. The signals used are shown in the tooltip, so the inference stays checkable.