A study of query fan-out in ChatGPT using brand searches

25 de August de 2026
26/08/202617:24

Every time someone asks ChatGPT about a brand, the system doesn’t search for that brand. It searches for a handful of other things, in parallel, that you as a user never get to see.

Most of what has been published about query fan-out looks at generic prompts — best CRM, how to choose a running shoe. I wanted to look at the scenario my clients actually worry about: what the model searches for when the prompt is about them.

So I ran a battery of 189 brand prompts across three of my clients, in three different sectors. That produced 1,797 searches performed by ChatGPT, none of them written by a human. This is what they look like, in what order they happen, and what we can take from them.

What I analysed

This is the methodology behind the numbers.

MetricValue
Projects analysed3 (healthcare, sports retail, hotels)
Brand prompts launched189
Prompts that produced fan-out174
Total runs723
Runs with fan-out687
Fan-out sub-queries1797
ModelGPT-5.5 (August 2026)
Prompt type100% branded, 12 categories

There’s an important nuance here: not every prompt triggers fan-out. In these three projects, 15 of the 189 prompts produced none at all, and 36 of the 723 runs resolved without a single search.

That means that in roughly 1 out of every 20 cases, the model already knew the answer and didn’t need to look anything up.

Everything else is what I’d call invisible demand: sub-queries that don’t exist in Search Console, that nobody types, and that nonetheless decide which sources make it into ChatGPT’s final answer.

Why I ran each prompt more than once

Fan-out is probabilistic. If you run a prompt a single time you get fewer than 3 sub-queries on average, but when you take it to 4 runs the number of distinct sub-queries climbs to 10.3.

In other words: a single run shows you roughly a quarter of what the model might ask about a brand, and you have no way of knowing which quarter you got. So I chose not to run each prompt once — to avoid noise, to avoid trusting purely circumstantial results, and to get closer to the structural patterns.

If you’re wondering how many repetitions your own monitoring needs in order to produce solid data, I’d recommend Gumshoe’s analysis of how much data you need to measure AI visibility with confidence.

Finally, the themes, families and groupings were consolidated by combining embeddings, tf-idf and detection of site: operators, plus regular expressions to identify brand names for the calculations.

As you’d expect, the goal here is to share patterns, so no client names are disclosed.

The headline numbers

A brand prompt doesn’t generate one search. It generates more than ten — and three out of every ten of those are aimed at a specific domain via the site: operator rather than at an open SERP.

MetricOverallRange across projects
Sub-queries per prompt10.39.7 – 10.7
Sub-queries per run2.6 (max. 13)2.5 – 2.7
Average sub-query length7 words6.7 – 8.0
Sub-queries containing the brand name93.4%92.6% – 94.5%
Sub-queries with site:30.2%25.9% – 37.4%

Only 7.8% of sub-queries are three words or shorter, and 15.8% are ten words or longer. That makes me think fan-out isn’t a “conversational long tail” at all — it looks more like there’s an SEO working backstage, writing these sub-queries out of the main prompt.

Prompts may well stay purely conversational, but the infrastructure underneath — LLMs that (for now) lack their own indexes — is still anchored to classic search engines and the way they work.

Three axes in branded fan-out

Before getting into the patterns, one thing to keep in mind: the same sub-query can carry a site: operator, quoted terms and a year all at once. So rather than one taxonomy, there are three independent cuts of the same dataset: where it searches, how it asks, and what it asks about.

Axis 1: where it searches

Destinationn%
Open SERP (no site:)125469.8%
site: [own domain]29916.6%
site: [group domain]1427.9%
site: [third party]1005.6%
site: [competitor]20.1%

Which means 1 in 6 sub-queries from a brand prompt goes specifically to your domain, and 7 in 10 carry no operator at all.

Axis 2: how it asks

Patternn%
site: operator54330.2%
Exact-match quotes42723.8%
Explicit year35219.6%
Booleans OR / parentheses241.3%
filetype: or PDF100.6%
Literal URL60.3%
Exclusion with -50.3%
inurl: / intitle:00%

This tells us it’s still crucial to have a project under control not just in terms of indexing, but in terms of being explicit in how content, headings, titles and descriptions are written.

Axis 3: what it asks about

Theme%
Generic brand23.3%
Product and service20.6%
Reputation and trust10.1%
Brand identity6.8%
Officialness and channels6.2%
ESG and sustainability5.2%
After-sales and support4.8%
Legal and compliance4.1%
Industrial property4.0%
Certifications and awards3.8%
People and leadership3.6%
Corporate reporting3.5%
Security and breaches2.3%
Press and media1.8%

With the three axes on the table, one number stands out: only 0.8% of sub-queries name a competitor. Branded fan-out is centripetal — it revolves around the entity itself, its group or satellite digital assets, and the sources that validate it.

That’s the big takeaway for any project: controlling your brand, your narrative and your digital footprint is the foundation of AI visibility.

The 7 patterns of branded fan-out

Pattern 1: own-domain fan-out

55.1% of all site: operators point at the brand’s main domain. This holds in all three projects, without exception.

site:[OWN DOMAIN] [corporate attribute]
site:[OWN DOMAIN] [product/service] [conditions]
site:[OWN DOMAIN] "[literal phrase]" "[literal phrase]"
site:[OWN DOMAIN]/[subfolder] [attribute]
site:[OWN DOMAIN] legal notice [BRAND] company number

This means the model is using your domain as a primary source of truth, querying it the way you’d query Google: give me everything you have on this domain.

So if you have content that isn’t indexed, or that can’t answer this kind of pattern, you’re leaving the door open to third parties.

One more figure: 64.4% of brand prompts trigger at least one site: search on the brand’s own domain. For the remaining 35.6%, the hypothesis is either that the model already had the answer, or that it didn’t consider the domain a useful source.

Pattern 2: group or parent-entity fan-out

26.2% of site: operators point at domains within the same group or parent: corporate sites, press rooms, sustainability or support subdomains, international assets.

site:[GROUP DOMAIN] annual report [year] [BRAND] [metric]
site:[GROUP DOMAIN] non-financial statement [year] [BRAND] [complaints|suppliers]
site:[GROUP DOMAIN] [BRAND] leadership people sustainability
site:[GROUP SUBDOMAIN] (sustainability|press|support) [attribute]

This has several implications, because any site within the group that talks about the brand becomes architecturally relevant for GEO. Depending on your global publishing strategy, citations could shift from one asset to another within the same group — and global domains may end up outranking local ones.

Pattern 3: third-party fan-out and external validation

18.4% of site: operators go to third parties. They’re the least numerous but among the most revealing, because they show who the model treats as an arbiter:

  • Official and industrial property registries: patent and trademark offices, patent databases, data protection authorities, company registries, tax ID lookups.
  • Social media, via official profiles.
  • Review and consumer platforms: opinion aggregators, complaint portals, consumer organisations.
  • Media and business press.
  • Sector databases and observatories.

The main implication is to widen the scope of the project and look at where you could improve your presence — even in places you don’t control today.

Pattern 4: literal fan-out (the quotes)

23.8% of sub-queries carry quotation marks, more heavily in the healthcare and hotel projects than in sports retail. Breaking down the 901 quoted strings:

Type of stringn%
2-4 words: brands, job titles, sections49955.4%
Single terms (1 word)30033.3%
Codes or IDs: tax numbers, case files, PDFs788.7%
Long phrases (5+ words)242.7%
"[literal phrase from your site]" "[BRAND]"
"[First Last]" "[BRAND]" LinkedIn
"[case file reference]" [authority] [incident]
"[file-name.pdf]" "[accounting line item]"

The model is quoting phrases it has already seen in order to confirm them. It’s a form of fact-checking against your own content. Claims, job titles and figures published on your site turn into facts to be verified.

Which is worth bearing in mind before taking creative liberties with copywriting, claims or CTAs: variation that reads well to a human can work against you in these new discovery channels.

Pattern 5: temporal fan-out

Freshness has been discussed at length in English-language studies. In mine, 19.6% of sub-queries contain an explicit year, with the figure varying widely by sector.

[BRAND] [country] official page returns warranties secure payment
[BRAND] core values
[BRAND] official slogan
[BRAND] parent company

Content that in another era was considered useless “for ranking” — and was sometimes even set to noindex — is no longer filler. These URLs can now build brand visibility, shield mentions and citations, and close the door on third parties.

We’re talking about pages like about us, legal notice, mission, vision and values.

Pattern 7: reputational risk fan-out

10.1% of fan-out sub-queries are about reputation and trust: opinions, reviews, ratings, complaints and grievances.

[BRAND] opinions trust complaints [review platform] [year]
[BRAND] reviews common problems
[BRAND] [incident type] complaints coverage
"[OWN DOMAIN]" "breach"
site:[THIRD-PARTY DOMAIN] [BRAND] breach

Fan-out can turn into an unsolicited reputation audit, which makes it vital to keep control of the narrative around incidents and complaints.

Ambiguous digital identity

This isn’t pattern number 8, but it’s worth calling out because it captures how LLMs affect brands rather well.

Going through every domain that appears in the site: operators, I found two things worth sharing.

It searches domains that are already redirected

In one of the projects, fan-out searches use site: on an old domain that already redirects to the current one.

Its prior knowledge was probably formed with the previous version of the site rather than the current one. And unlike the temporal pattern, where it actively verifies freshness, here it doesn’t question domains at all: it takes them at face value.

That has a direct implication for how we handle a project’s infrastructure and its map of digital assets: be tidy, and keep basic SEO hygiene with a proper redirect inventory.

It tries several social profiles for the same brand

Social media accounts for 5.5% of fan-out sub-queries, and they don’t always go to a specific profile.

In one project the model tried three different Instagram accounts, two LinkedIn, two Facebook and one YouTube, looking for the brand’s social profiles. Through that trial and error it’s trying to guess which ones are correct.

This connects directly with pattern 6 and the whole officialness theme: if it’s searching for that, it’s because it hasn’t resolved the information.

The fix is unglamorous: clear communication, links from your site, schema markup declaring your profiles, consistent naming across accounts. It sounds basic — and that’s the point. Leaving the basics unresolved is what slowly costs you control of your brand inside LLMs.

The multilingual insight

Worth watching closely: 8.9% of sub-queries use English vocabulary. In projects with an international footprint that rises to 27.3%, and in a purely local project it drops to 1.0%. On top of that, 9.9% of runs mix languages.

So if your brand has an international presence, the model is going to bring English into play even when users prompt in Spanish, and your global assets will show up. Which makes consistency of messaging across language versions matter a great deal — not translation as such, but which content each asset actually tells.

A finding about how fan-out escalates

Now let’s try to understand the order in which ChatGPT searches when we give it brand prompts, because the resulting fan-outs follow a sequence — and there are signals to read in it.

It starts by asking and ends by checking

From the 1,797 sub-queries we can sort by the position each one occupies within its run, and a pattern emerges:

  • The first search is natural language.
  • From the second onwards it changes: the site: operator appears, narrowing the search to a specific domain.
  • After that, quotation marks start showing up, specifying exact phrases, word for word.

The chart shows exactly that: first it narrows, then it verifies. At the start it needs to know where to look; after that it needs to check whether the source says precisely what it thinks — and the second part only happens when the first wasn’t enough.

You can also see the queries getting shorter, as the model stops describing.

First versus last search

In fact, if we compare only the first search of each run with the last, it’s hard to argue with:

Use of site: multiplies by 4. Quotes multiply by 25.

The sequence, step by step

Putting it all together, the journey these fan-outs took across three sectors can be summed up in four phases:

  1. Open question: natural language, no operators, no quotes.
  2. Narrow: the site: operator.
  3. Verify: quotes appear, checking whether something is stated literally.
  4. Insist: shorter, more specific reformulations.

Here’s a real anonymised example, using brackets to show the patterns. The prompt asked which events the brand takes part in.

It looks at the brand’s site, doesn’t find what it needs, searches quoted phrases, goes to a third-party site, comes back, and ends up searching a literal sentence from a press release.

Those 11 steps hint at the opportunity in working on organic and AI visibility together, keeping in mind how both search engines and models actually operate.

Escalating is not the norm

That said, most runs don’t escalate. What does that mean?

Of the 687 runs with fan-out, 246 make a single search and 193 make two: 63.9% are resolved in two searches or fewer — on top of the runs that produce no fan-out at all.

So escalating isn’t the default behaviour. It’s what happens when everything before it hasn’t worked.

There’s more. Sometimes the model asks the same thing several different ways: in 23.1% of runs with two or more searches there’s at least one pair of sub-queries that are identical or barely differ by a few terms — sometimes just the brand name on its own, its values, mission or slogan.

Which leaves another useful insight: if the model insists on making similar or identical searches and you’re not in one of them, you’re probably not in any of them.

This study will need repeating and widening to more brands, but there’s already a hypothesis on the table: a long fan-out may mean the model is rummaging around, and that opens the door to ending up on third-party sites instead of the brand’s own.

Actionable takeaways

  1. One indexable URL per corporate attribute.
  2. Key data in literal HTML, not only in PDFs or inside images.
  3. Consistency in claims and slogans, so they hold up under exact-match searches.
  4. Visible update dates and current years in corporate and other content.
  5. Your own pages for reviews, complaints and incidents.
  6. Coherence between each market’s main domain and the group domains.
  7. Single, linked official social profiles.
  8. Correct presence and data in official registries.
  9. Live redirects from past migrations.

Limitations of this study

  • Three projects, three sectors. The patterns repeat across all three, but three is not a representative sample of the market.
  • One model, one moment. GPT-5.5, August 2026, just before the GPT-5.6 release.
  • One capture method. Data was extracted through OpenAI’s official API, which may differ from what real users see.
  • Brand prompts only. None of this extrapolates to generic prompts.
  • We observe what was retrieved, not what was cited. This is about what the model searched for, not about the final answer, its mentions or its citations.
  • Observations within the same run are not independent, so real margins of error are somewhat larger than the calculated ones.

References

Soy MJ Cachón

Consultora SEO desde 2008, directora de la agencia SEO Laika. Volcada en unir el análisis de datos y el SEO estratégico, con business intelligence usando R, Screaming Frog, SISTRIX, Sitebulb y otras fuentes de datos. Mi filosofía: aprender y compartir.

Leave a Reply

Your email address will not be published. Required fields are marked *