When Google shows a user an AI Overview, it provides an answer that cites a handful of pages as sources.
The question that instinctively comes to mind: between two pages competing for the same search, why does it cite one and not the other? Does it take into account “how it’s written” at a linguistic level?
There are plenty of rumours and assumptions out there: better to write short paragraphs, better to put definitions at the top, better to use lists…
To dig into this, I ran an express study of 500 search queries in the Spanish sports sector. I collected the answers, downloaded the content of around 2,000 pages and measured more than 50 linguistic features at paragraph level (over 75,000 paragraphs in total).
Who does Google cite in its AI Overviews? The short answer is that it cites, above all, whoever ranks well in the organic results, and also those whose content answers the query well (although this matters less than I expected).
The finding I didn’t expect (not even remotely) is that more than 30% of the answers end by inviting and encouraging the user to keep talking, asking and interacting, with sentences that Google’s systems write without citing anyone, simply inviting the user to leave the SERP and continue in AI Mode.
There are a few more things I’ll tell you about.
The study in one minute
- Organic position is the most notable and consistent factor. A page in position 1 is cited in 60% of AI Overviews; a page in position 5, in 34%.
- Domain size, measured by its organic traffic, does not predict citation.
- Answering the question helps, but less than expected.
- There is a slight advantage for more developed, complete pages, with more paragraphs and words.
- The shape of the text (closed paragraphs, varied vocabulary) has small effects on citation.
- Lexical richness is the writing feature with the most weight, although its effect is small.
- YouTube is the most present source: it is cited in 77% of AI Overviews, with a median of two videos per answer.
- Google rewrites from its sources, but doesn’t copy them word for word. 84% of the sentences in the answers don’t keep the wording of the sources.
- 30% of AI Overviews open a conversation in their own words and “hijack” the user, taking them from the SERP to AI Mode.
What this analysis consists of
I took the 500 queries from SISTRIX, using one of the sector’s leaders and its top-100 rankings. To balance the sample, I didn’t use the most searched queries: instead, there are mini-samples by position and volume bands, so that a bit of everything is represented.
| Metric | Value |
|---|---|
| Queries | 500 |
| With AI Overviews | 451 (90.2%) |
| Distinct answers | 417 (92.5%) |
| Total citations in AI Overviews | 2,694 (2,209 distinct URLs, 854 domains) |
| Of which, YouTube videos | 857 (719 distinct videos) |
| Web pages analysed | 1,936 (1,092 cited, 780 control, and others outside the ranking or without a pair) |
| Paragraphs extracted and measured | 79,060 |
| Linguistic indicators per paragraph | 53 |
| AI Overview sentences analysed | 4,447 |
| Comparisons in the main model | 1,913 across 377 searches |
| Capture and extraction | DataForSEO API, 17 September 2026 |
| Market | Spain, Spanish, desktop, logged-out session |
One important detail: whenever I compare what gets cited with what doesn’t, it’s always done within the same search, without mixing, in order to isolate the factors.
How the dataset was built
- SERP and AI Overview capture: using the DataForSEO API, for each query I stored the answer text block by block, the sources for each block and the organic ranking.
- Content download: for every cited URL and for up to 3 non-cited organic pages per search, which make up the control group. For YouTube videos, I downloaded each video’s transcript.
- Paragraph segmentation and calculation of 53 linguistic indicators.
- Page-level aggregation: the median of each indicator, number of paragraphs, total words, share of paragraphs ending in a full stop, age and domain traffic.
- Search-level variables: the organic position of the URL in that search and the semantic similarity between the page and the query.
- Comparison of cited versus non-cited pages within each search: using a conditional logistic regression model.
What each linguistic indicator measures
| Family | Examples |
|---|---|
| Length | words, sentences, characters, average sentence length |
| Parts of speech | density of nouns, verbs, adjectives, adverbs, pronouns, proper nouns |
| Syntactic complexity | dependency tree depth, subordination, relative clauses, coordination |
| Voice | periphrastic passive, reflexive passive, starting with a verb |
| Concrete information | entities, figures, units of measurement |
| Paragraph autonomy | anaphora without an antecedent, deixis (“as we saw above”), definitional pattern (“X is…”) |
| Modality | hedging, intensification, obligation, imperatives |
| Person | first and second person, questions |
| Connectors | causal, adversative, additive, reformulative, exemplifying, sequencing |
| Form | starts with a capital letter, ends with a full stop, brackets, “concept: explanation” structure |
| Readability | Fernández Huerta, Szigriszt, syllables per word, long words |
| Relevance | semantic similarity between the text and the query (multilingual embeddings) |
What an AI Overview looks like in the sports sector
An AI Overview is a search feature that generates an answer to the user’s query using artificial intelligence.
On desktop it typically appears at the very top of the results page, although it can also show up in less prominent positions. On mobile, it usually takes up much of the top of the screen.
This is how it looks on desktop:

These are the study’s key figures about the feature itself:
| Metric | Value |
|---|---|
| Queries with an AI Overview | 90.2% |
| Sources cited per AI Overview (median) | 6 (mean 6.5; maximum 17) |
| Sentences per AI Overview (median) | 10 (between 4 and 30) |
| Headed sections per AI Overview (median) | 2 |
| Textually unique answers | 92.5% |
| Blocks with no declared source | 28% |
Of the 500 keywords reviewed, more than 90% triggered AI Overviews, most likely for a simple reason: they are informational sports and health queries, which is where Google can potentially deploy the most AI Overviews.
Next, it’s worth talking about the 92.5% of textually unique answers. To calculate this, a digital fingerprint or hash is created from the content of each AI Overview. That alphanumeric code will only be the same when two AI Overview answers are identical: it’s binary, either it matches or it doesn’t.
Methodologically, this was important so as not to count duplicated cases as independent texts. If I had counted them, I would have been giving importance to things that don’t really have it.
But if we look at the remaining 7.5%, we find something interesting: there are 20 groups of queries that share the same answer.
Example: “empezar a correr”, “quiero empezar a correr”, “correr como empezar”, “como iniciarse a correr”, “iniciarse a correr”, “quiero aprender a correr” (all variations of “how to start running”).
To check that this wasn’t an error in the API or in the analysis, I ran two checks two days after the API extraction:
- The first, by hand, using the 6 example queries, in incognito mode and logged out. 2 of them still returned the same answer, text and sources; the other 4 didn’t.


- The second, through the same API. I relaunched the 54 queries from the 20 groups identified, using the same method as the study. The result: no identical answers, the number of sources stayed the same but only half of them were the same, and 1 query no longer showed an AI Overview. However, 15 of the 20 groups were still sharing an answer, even though that answer was different from the one two days earlier.
In the example group, five of the six queries were still sharing an answer, and all five had lost the same domain and gained the same new one. This suggests that when Google changes a source, it changes it for every query in the group at the same time.
On the other hand, the fact that the browser and the API don’t match on the same day isn’t an error either. It’s one more layer of variability: what a person sees with their location and history doesn’t have to be the same as what an API query returns.
The two conclusions:
- There is a grouping of queries by intent, and it’s fairly stable. It holds in 75% of the groups two days later. Queries in the same group share sources, but if Google changes one, it changes them all. So if you appear, you appear in all of them, but if you disappear, you disappear from all of them too.
- Google doesn’t keep the same content or sources in AI Overview answers from one day to the next. In 48 hours, every answer had changed and half of the cited sources were different. This points to volatility.
That’s why the big insight is that AI Overview monitoring can’t be limited to a single day: what needs to be measured is frequency of appearance over time.
Who gets cited
Four of the five most present domains are social or video platforms, such as Instagram, TikTok and Facebook, which appear in more AI Overviews than Wikipedia.
We can say that, for informational sports queries, Google is building answers with creator content, not just with media outlets and brands.
| Domain | AI Overviews it appears in | % of all AI Overviews |
|---|---|---|
| youtube.com | 321 | 77% |
| instagram.com | 77 | 18.5% |
| tiktok.com | 65 | 15.6% |
| facebook.com | 62 | 14.9% |
| es.wikipedia.org | 52 | 12.5% |
| reddit.com | 26 | 6.2% |
| redbull.com | 25 | 6% |
| vivagym.com | 23 | 5.5% |
| nike.com | 22 | 5.3% |
| mayoclinic.org | 20 | 4.8% |
Below these “social” leaders, the distribution is very flat (seven out of ten domains appear in just one AI Overview), and appearances are not concentrated in a group of favourite sources.

The weight of YouTube
There are several things to bear in mind here:
- YouTube is cited in 77% of AI Overviews, with a median of 2 videos per answer and a maximum of 11. In total there are 719 distinct videos, spread across hundreds of channels.
- Videos are cited within the text, not just in the side panel. 851 of the 857 cited videos appear as the source of a specific block of the answer, not only in the side panel.
- Almost none of the cited videos appear in the classic organic results for that search: only 7%.
- No AI Overview is built from videos alone. Video accompanies text; it never replaces it.

Organic position is still the way in
| Organic position | Cited in the AIO |
|---|---|
| 1 | 60.4% |
| 2 | 49.2% |
| 3 | 43.6% |
| 4 | 40.3% |
| 5 | 34.1% |
| 6 | 31.9% |
| 7 | 32.2% |
| 8 | 28.9% |
| 9 | 22.4% |
- Position one doubles the chances of position eight.
- The biggest jump is between the first and second positions; from there on, the drop is gradual.
- “Position” here refers to position among the organic results, not counting the AI Overview itself, videos or other SERP modules or blocks.
What predicts being cited
To find out what makes a page get cited or not, I ran more than 1,900 comparisons across 377 searches. These are the main factors with an effect:
| Factor | Effect |
|---|---|
| Lexical richness | ×1.22 |
| Better organic position | ×1.20 |
| Similarity to the query | ×1.18 |
| Number of paragraphs | ×1.16 |
| Paragraphs ending in a full stop | ×1.15 |
How to interpret this:
- No single factor dominates. All of them are similar and small: there isn’t “one thing” that gets you cited on its own; it’s the sum of several things.
- What each one means:
- Lexical richness: texts with varied vocabulary that don’t keep repeating the same terms.
- Organic position: the higher up, the better the chances.
- Similarity to the query: covering the same topic the user is asking about, not necessarily with the same words. It’s not the biggest effect, but it adds up.
- More paragraphs: more developed pages have a slight advantage.
- Paragraphs ending in a full stop: this refers to complete sentences.
- Other aspects that don’t add up: paragraph length, using more nouns or more verbs, starting with definitions…
- Domain traffic predicts nothing, so being a big site doesn’t get you cited. What counts is the specific page and its position for that search.
- Freshness is a hint, not a conclusion. Cited pages were updated more recently than non-cited ones, but it’s a weak signal. Given that some pages don’t declare a date, it’s harder to analyse and to include among the factors.
Does anything change when you split the data into segments, for example by type of search? I checked, and found no differences. So there isn’t a specific recipe for video-heavy queries and another for text queries.
Google doesn’t copy, it rewrites
Another thing I did was compare each AI Overview sentence against the text of the pages it cites as sources, looking for identical word sequences, to answer this question: does Google rewrite, copy, or what does it do in these AI-generated texts?
| What Google does with the source | % of sentences |
|---|---|
| Rewrites | 84% |
| Paraphrases | 14% |
| Copies word for word (or almost) | 2% |
So only 2 out of every 100 sentences keep the source’s wording. The most common thing is for Google to say it in other words.
One in three AI Overviews invites the user to leave the results page
It doesn’t give a closed answer but an open one. This means that in the final sentences of the answer, it adds something that “invites” or “nudges” the user to keep the conversation going, for example:
- “Si me cuentas cuál es tu nivel actual y qué dificultad se te resiste más, te puedo dar recomendaciones más adaptadas a ti.” (“If you tell me your current level and what you’re struggling with most, I can give you recommendations better suited to you.”)
- “Te puedo sugerir una distribución de rutina semanal adaptada a tu tiempo.” (“I can suggest a weekly routine plan adapted to your schedule.”)

This happens in 30% of the answers (127 of 417), and 126 of those 127 sentences don’t cite any source. Google writes them, with the intention of “stretching” or prolonging the conversation. Broadly speaking, they follow three patterns:
- Disambiguating. The aim is to clarify what you want. One example found in the data was the search “descompresión” (decompression), where the final sentence was “Dime si te interesa el buceo, la medicina o la informática” (“Tell me whether you’re interested in diving, medicine or computing”).
- Asking for context. The aim is to find out who you are or what your habits are, that is, to learn more about you in order to tailor the answer. For example, for the search “adelgazar nadando” (losing weight by swimming), the closing sentence was “¿Cuántos días a la semana puedes ir a la piscina?” (“How many days a week can you go to the pool?”).
- Doing something for you. The aim, or the proactivity, is to offer to do something for you, as in the example search “core ejercicios” (core exercises), whose final sentence was “Te puedo preparar una rutina personalizada…” (“I can put together a personalised routine for you…”).
And what happens next? If the user plays along with Google, the conversation takes place outside the results page, although it doesn’t disappear completely: it leaves a trace in Search Console.
Anastasia Kourou already anticipated this and John Mueller confirmed it. I’m analysing the data and will cover it in a separate article, with the queries worth filtering and what to do with them.
Why only Google?
The truth is that I tried to replicate the study on Bing too, and it wasn’t possible using DataForSEO because, according to their support team, for many queries the Copilot answer isn’t included in the page’s HTML: the browser generates it afterwards, using scripts.
So with their API, there’s no way to extract it. In future studies I’ll look for alternatives, and if you, reading this article, know of one, leave a comment.
What you can do with this in your project
- Prioritise position. Position one is cited in 60% of cases; position five, in 34%.
- Answer what’s being asked. You don’t need to repeat the keyword, but you do need to cover what the user is looking for (or expects to find).
- Write with varied vocabulary and in complete sentences. Lexical richness and closed paragraphs are the writing features with the most weight, although their effect is still small.
- Don’t look for writing shortcuts. No single factor gets you cited on its own, so be wary of checklists, tricks and shortcuts on “how to write for AI Overviews”.
- Keep an eye on video and social platforms in your sector. YouTube appears in 77% of answers, and Instagram, TikTok and Facebook in more than Wikipedia. Social search is a cross-cutting reality, more than ever.
- Measure frequency, not presence. In two days, Google rewrote every answer and changed half of the sources, so appearing on a specific day doesn’t mean much.
Limitations of this study
- One sector, one language, one market: sport, Spanish, Spain.
- The sample isn’t random: the keywords come from the rankings of a single site in the sector.
- A “snapshot” of a single moment: the answers were captured on 17 September 2026, via API and logged out; what each person sees may be different.
- Google only: it wasn’t possible to reliably extract Copilot answers on Bing.
- Format not analysed: I can’t say whether lists or tables are cited more than plain text.
- It’s observational: it shows what distinguishes cited pages, not that changing something will get you cited.
- Keep your content up to date. The signal isn’t conclusive, but it points in the expected direction.
How long it takes to produce this kind of content
About two weeks:
- Defining what to study
- Choosing the extraction, processing, cleaning and consolidation methods
- Testing the data and segmenting statistical models
- Writing the article with the main data and insights, in Spanish and English
In terms of cost, DataForSEO’s APIs for extracting Google results pages are not expensive; the ones for content extraction and parsing, and for domain metrics, are a different story.
Soy MJ Cachón
Consultora SEO desde 2008, directora de la agencia SEO Laika. Volcada en unir el análisis de datos y el SEO estratégico, con business intelligence usando R, Screaming Frog, SISTRIX, Sitebulb y otras fuentes de datos. Mi filosofía: aprender y compartir.
