Ten or more hostnames cited in 28 of 253 sampled Google AI Overviews responses
About one in nine sampled Google AI Overviews responses with citations reached ten distinct hostnames in Bee LLM Observatory data. That describes the breadth of references, not their quality.
An extensive source list in a minority of responses
Of 253 observed Google AI Overviews responses with citations, 28 cited at least ten distinct hostnames: 11.07%, or approximately one in nine. The denominator matters. This percentage applies only to sampled responses containing valid citations, and does not show how often all responses from that channel include such an extensive set of references.
These were questions monitored by Bee LLM, with unequal weights across questions and accounts. The result describes this particular sample rather than the typical experience of every user. Within that scope, responses reaching such a broad list of cited hosts were a minority.
| Engine / channel | Responses with citations | Responses meeting the criterion | Percentage |
|---|---|---|---|
| Gemini | 400 | 25 | 6.25% |
| Google AI Overviews | 253 | 28 | 11.07% |
What reaching ten actually means
The measure counts distinct hostnames in a response’s citations. A hostname is the part of a web address identifying the server or site it points to. Several cited pages on the same hostname do not multiply its contribution to the count: reaching this threshold depends on the number of different names.
That distinction matters because ten hostnames do not necessarily represent ten independent organizations. Names can include subdomains, and the count does not investigate who controls each source. Nor does it establish that a response contains ten separate arguments, ten checks of a claim or ten contrasting perspectives. It measures the extent of the cited host list.
More references do not establish greater reliability
The most direct reading is that some responses assemble references spread across quite a few web destinations. That may interest readers who want to explore the sources. However, the aggregate cannot reveal whether each link adds information, repeats another reference or adequately supports the statement it accompanies.
The threshold also treats a response citing exactly ten hostnames the same as one citing many more. Without examining the content and the connection between claims and sources, this percentage cannot serve as a quality score. A response below the threshold could be well documented, while one above it could still need additional checking.
What happens when every question has equal weight
The primary rate, 11.07%, gives every response equal weight. A question run more frequently therefore contributes more observations to the percentage. The Google AI Overviews segment covers 70 prompts: the questions or instructions used to obtain responses.
A sensitivity calculation gives each prompt equal weight instead, producing a rate of 10.52%. The interpretation remains similar: roughly a tenth reach the threshold. This alternative is neither a margin of error nor an estimate for the general population. It shows how the summary changes when questions receive different weights; it does not also give accounts equal weight.
Another observed channel, without a controlled comparison
The published aggregate also includes Gemini. There, 25 of 400 responses with citations reached ten distinct hostnames, or 6.25%, across 101 prompts. Giving prompts equal weight produces 7.78%. For this segment, the adjustment raises the percentage, although responses meeting the threshold remain a minority.
Both channels’ figures are descriptive. They do not come from a test using identical questions and conditions: prompts, account weights and collection methods may differ. They cannot support an engine ranking. Engine labels identify collection channels, and an API channel should not be equated with its corresponding consumer application. Engines absent from the aggregate cannot be assigned a zero.
When and how the sample was collected
This article is published on September 29, 2026. The observations cover seven earlier, complete days: from September 21 at 00:00 up to, but excluding, September 28 at 00:00, in the Europe/Madrid time zone. The publication date falls outside the observation period.
Selection retains the latest response for each execution and stored model, using only completed executions with nonempty text that is not an error. Branded prompts and accounts with “demo” in their email are excluded. The measure uses recorded citations with a known collection method, rather than any web address detected in the text. Empty or missing citation records are excluded from the headline denominator.
The coverage behind the finding
Quality checks record 2,459 valid responses from known engines. Citation telemetry was missing for 18, leaving 2,441 measured responses; 2,138 contained citations. That citation-bearing set covers 161 prompts, 29 projects and 26 accounts. These totals describe the aggregate’s overall coverage, rather than the size of each published segment.
Publication requires segments to contain at least 100 responses, ten prompts, three projects and three accounts; responses meeting the measure must span at least three accounts. These minimums provide coverage, but do not make the sample representative. The data describes neither all AI conversations nor market share, and publishes no individual information about people, accounts, projects, prompts or responses.
Frequently asked questions
Does 11.07% refer to all responses?
No. It refers to the 253 sampled Google AI Overviews responses with valid citations. Of those, 28 reached the threshold.
Do ten hostnames mean ten independent sources?
Not necessarily. The count identifies different names in cited addresses, but does not verify that they belong to independent organizations.
Does publishing on more sites guarantee citations?
No. This aggregate describes observed references and establishes no causal relationship between publishing content and receiving citations.
