For two months, a technical oversight left a door open. Between 21 May and 21 July 2026, the data ChatGPT sends to your browser contained a field naming, for every result, the source it came from. The field has since disappeared. While it was visible, it proved something the industry suspected but could not demonstrate: ChatGPT no longer just queries Bing — it owns an index of the web, called Labrador. For a local business, the consequence is not the one you would expect, and it fits in two hundred characters.
What the leftover field revealed
The field was called result_source, and it only ever carried four values: labrador, bright, oxylabs and serp. The last three are providers that fetch Google results on OpenAI's behalf. The first, documented in September 2026 by Tomek Rudzki, is OpenAI's own index.
And it is not one index, it is a family. The observations describe a dozen specialised indexes: general web, news (with separate tiers for the last day, the last week, the archive), PDF, YouTube, Wikipedia, arXiv, finance, legal, medical, shopping, images — and a local index. That last one is ours.
OpenAI's stated internal ambition, as far back as 2023, was to answer 80 % of queries from its own index alone. It is nowhere near: the analyses agree that full independence will take years. But the direction is set, and it invalidates a sentence I have written on this blog myself, and that you will read everywhere else: "ChatGPT searches on Bing". It was true. It is no longer the whole truth.
What I'm correcting in my own articles. In "why your business doesn't show up in ChatGPT", I explained that working on Bing was the way in. That is still a sound reflex, and a useful floor test: a page absent from Bing and absent from OpenAI's crawl has no chance. But it is no longer sufficient, because OpenAI now crawls the web itself — and that part you control from your own site.
OpenAI's three crawlers, and the one that decides your visibility
This is the most misunderstood point, and the most expensive one. OpenAI runs three distinct crawlers with three different jobs. Blocking or allowing them does not have remotely the same consequences.
| Crawler | What it does | If you block it |
|---|---|---|
| OAI-SearchBot | Builds and refreshes ChatGPT's search index | You disappear from answers |
| ChatGPT-User | Opens a page live during a conversation | The model can no longer read your page in detail |
| GPTBot | Collects content for model training | No effect on your search visibility |
Plenty of business owners have, without knowing it, exactly the opposite of what they want. They read that they should "protect themselves from AI", or a security plugin, a purchased theme or a CDN firewall added a generic "block AI bots" rule. The result: invisible in ChatGPT, without having protected anything worth protecting, since training and search are two separate things.
The right setup for a business that wants to be recommended is simple: OAI-SearchBot allowed, ChatGPT-User allowed, and GPTBot according to your stance on training — allowing it gains you no visibility, blocking it costs you none. That is the configuration this site runs, and you can check your own in seconds: type yourdomain.com/robots.txt into your browser. If you find Disallow under OAI-SearchBot, or a rule blocking everything, that is today's priority.
The finding that should change how you write
Knowing you are crawled says nothing about whether you are read. The most useful study on that is French. In July 2026, RESONEO analysed how ChatGPT's retrieval actually behaves across a substantial corpus: 1,249 real conversations, 88,000 search results, 26,900 distinct pages, on free and paid accounts in several countries.
| What the study measures | Result |
|---|---|
| Answers that open no page at all | 9 in 10 |
| Retrieved pages actually opened | 1 in 80 |
| An opened page ends up cited | 74 % of the time (440 pages) |
| A merely listed page ends up cited | 7 % of the time (57,853 pages) |
| Text considered without opening | title + ~200 characters from the H1 |
Read the last two rows together, because that is where the lesson sits. Most of the time, ChatGPT does not read your page — it reads a thumbnail. Your title, then roughly two hundred characters starting at your main heading. Everything below that — your carefully written "about" page, your hours in the footer, your menu in a PDF — plays no part in the decision.
Two hundred characters is about the length of the paragraph you just read. It is very little, and it is very concrete: that is your entire budget for existing in nine answers out of ten.
What that looks like, page by page
The instinct on business websites is to open with a flourish: "Welcome", "Craftsmanship since 1987", "Your satisfaction is our priority". Top of the page, large type, usually over a photo. In the old world, that was harmless. In an index that keeps only two hundred characters, you have just spent your entire budget saying nothing.
Compare on a real case, a bakery in east Paris.
| What gets extracted | What the AI takes from it |
|---|---|
| "Welcome! For three generations, our family has upheld a tradition of artisanal craftsmanship…" | A bakery. Somewhere. Like 30,000 others. |
| "Artisan bakery on rue de la Roquette, Paris 11th. Natural sourdough breads, open Tuesday to Sunday 7am–8pm, closed Monday. Gluten-free on Thursdays. Celebration cakes to order, 48h notice." | A precise address, hours, two rare specifics: citable. |
The second version is not "better written". It is dense with verifiable facts, which is exactly what a model looks for when deciding whether to recommend you for "gluten-free bakery open Sunday near Bastille". The rule fits in one sentence: the first sentence after your title must answer the question, not introduce it. It is the same principle I set out for Google AI Mode's query fan-out, and that is no coincidence: all these systems break pages into standalone passages.
The thirty-second test. Open your home page. Copy your title, then the first two lines of actual text under your main heading — not the menu, not the banner, the text. Paste that into a document and read it back. If a stranger cannot work out what you sell, where, and when you are open, you have found this week's job. Repeat on your three most important pages.
The cache trap: a mistake can follow you for months
The RESONEO study describes a second mechanism, less known and more awkward. When ChatGPT does open a page, it keeps a copy. That copy then serves other conversations, including other users, without your server being contacted again. And it is kept for a long time: several months, with no observed expiry.
Translated for a local business: if ChatGPT read you during your annual closure, or with your old hours, or with a price you have since changed, that version can keep circulating well after you updated the page. You cannot force a refresh, which is one more argument for two simple habits: announce exceptional closures in advance rather than editing in a panic, and never let perishable information exist only inside an image or a PDF, where it is even harder to correct.
It is also why consistency between your site, your Google Business Profile and directories matters more than ever: when one source goes stale, the others arbitrate. I set out who feeds whom in the article on local data providers — in France, Pages Jaunes carries real weight in what ChatGPT knows about your street.
What this does not mean
I am not going to sell you "Labrador optimisation", and be wary of anyone who does. Three reasons.
- Nobody knows Labrador's real share. The available measurements show it varies enormously by mode: dominant on free accounts in instant answers, nearly absent in reasoning mode on paid accounts, where scraped Google results dominate instead. Optimising for one may change nothing for the other.
- The index is rebuilt continuously. What held in August 2026 will not necessarily hold in January. The very field that made this analysis possible has already gone.
- Other AI tools do not work this way. Gemini runs on Google's infrastructure; Perplexity and Claude have their own pipelines. A site tuned for one system is a fragile site.
What survives all of it is the groundwork, and there is nothing exotic about it: a site crawlers can read, titles that say what the page contains, an answer in the first sentence, hours and prices as text, and the same information everywhere. That groundwork serves Labrador, Google, Gemini and the hurried human on a phone, and it is the only investment that does not go obsolete when an index changes name.
The checklist, in order
- Check your robots.txt — type your domain followed by /robots.txt. OAI-SearchBot and ChatGPT-User must be allowed. If you see a rule blocking everything, or a Disallow inherited from a plugin, that is job number one.
- Run the thirty-second test on your three main pages: title plus first two lines. Trade, town and hours should all be there.
- Rewrite the first sentence of those pages so it answers instead of welcoming. Keep the flourish if you like it — put it after.
- Get perishable information out of images and PDFs. Hours, menu, prices: as text on the page. The cache remembers for a long time, so it may as well remember something accurate.
- Check consistency across site, Google listing and directories. If two sources disagree on your hours, the freshest one wins — and it may not be yours.
- Measure. Search Console cannot see ChatGPT, but visits coming from assistants are identifiable in your analytics; I explain how to isolate them in measuring your AI visibility.
Frequently asked questions
Does ChatGPT still use Bing?
Partly, and less and less. Bing supplied the original index and still contributes, but OpenAI has run its own index since 2026, fed by its OAI-SearchBot crawler, alongside Google results fetched through providers. Being indexed in Bing remains a useful floor test: a page absent from Bing and absent from OpenAI's crawl has no chance of being retrieved. But it is no longer the single door in, and the part you genuinely control now sits on your own site, in your robots.txt and in your opening lines.
What is the Labrador index?
It is OpenAI's in-house web index. Not a single index, but a family of specialised ones: general web, news, PDF, YouTube, Wikipedia, finance, legal, medical, shopping, images, and a local index. Its existence was documented in 2026 from the technical data ChatGPT sends to browsers, during the two months a field named it explicitly. Nobody outside OpenAI knows its exact share of answers, and that share varies with the version in use.
Should I allow OAI-SearchBot?
Yes, if you want to appear in ChatGPT. It is the crawler that builds the search index: blocked, your site cannot be retained as a source. It is distinct from GPTBot, which handles training, and the two are controlled separately — you can perfectly well allow the first and refuse the second. The classic trap is a "block AI" rule added by a plugin or firewall, which blocks both indiscriminately and makes you invisible while protecting nothing useful.
Why am I crawled but never cited?
Because being crawled, being retrieved and being opened are three different stages. The RESONEO analysis from July 2026 measures that nine answers in ten open no page, and that only one page in eighty is actually opened. An opened page is cited 74 % of the time, against 7 % for a page merely present in the list. So most of the time ChatGPT judges you on a short extract, which shifts the work from the page's content to its first two hundred useful characters.
Which part of my page is actually read?
Without the page being opened: the title, then roughly two hundred characters from the H1. That is all. For a local business, it means the trade, the town and the distinctive detail must sit there, not in the third paragraph or a welcome banner with no text. It is the highest-return change you can make this week, and it costs only the time to rewrite three sentences.
Should I rewrite my whole site?
No. Labrador is rebuilt continuously and the other assistants work differently: optimising for one index at one moment is chasing a moving target. The groundwork does not move — crawler access, explicit titles, an answer in the first sentence, hours and prices as text, consistency across your sources. That groundwork serves every engine, Google included, and it is exactly what I put in place in an AI search optimization engagement.
If you want to know where you stand, send me your site address: I'll check your robots.txt, extract what ChatGPT actually sees of your three main pages, and tell you in one page what to change. That is the starting point of my work on AI visibility, and the groundwork is the same one behind every business site I build: say what you do, where, and when — from the first line.