Skip to content
Original study

One site cited in 103 of 558 sampled OpenAI API responses with citations

Citations pointed to a single site in nearly one in five sampled responses with citations from this channel. How often individual prompts were repeated affects that finding.

5 min read

Roughly one in five, within this sample

In Bee LLM’s OpenAI API sample, 103 of 558 observed responses with citations cited exactly one distinct hostname: 18.46%. In everyday terms, their references pointed to a single site. This finding applies to sampled responses containing citations, rather than every response from the channel or AI conversations generally.

The channel’s sample covers 145 distinct prompts, meaning questions or instructions. Some were repeated more often than others, and accounts also carried unequal weight. That composition matters: giving every prompt equal weight reduces the proportion to 14.43%. OpenAI API identifies a collection channel here; it does not represent the consumer ChatGPT app.

Results for the sample described in this article
Engine / channelResponses with citationsResponses meeting the criterionPercentage
ChatGPT (web)318299.12%
OpenAI API55810318.46%
Gemini456102.19%
Copilot25941.54%
Google AI Mode1994120.60%

One site can supply several references

The measure counts distinct hostnames in each response’s citations. A hostname is the part of a web address that identifies the site being linked to. To qualify as one of the 103 responses in the finding, every valid citation in a response must point to that same hostname.

Citing one site therefore does not necessarily mean providing only one citation or linking to a single page. A response can reference several pages on the same site and still qualify. The count also does not measure authors, independent documents or media owners. It describes how concentrated the cited web addresses are, within that specific definition.

Repeated prompts affect the headline figure

The primary rate is 18.46% because each observed response has equal weight in the calculation: 103 qualifying cases out of 558 responses with citations. This describes what happened across the collected responses. But a prompt that runs many times contributes more results to the overall percentage than a prompt that runs less often.

The alternative calculation gives every prompt equal weight and produces 14.43%, a reduction of 4.03 percentage points. That shows that repetition patterns influence the main figure. Single-site citation remains a recurring feature under this weighting, but it is less frequent. The adjustment neither makes the sample representative nor gives every account equal weight.

Concentration does not establish quality

For someone reading an AI response, the finding offers a concrete observation: several visible references can lead back to the same site. Counting links and counting sites answer different questions. This measure describes the variety of cited destinations, but it cannot establish whether the references adequately support a response.

The aggregate evidence does not assess answer accuracy, the authority of the sites or whether separate pages provide independent evidence. It also cannot explain why a single site appeared in any particular case. Calling these 103 responses better, worse or biased would require information and evaluation that this count does not supply.

An observation window, not a universal picture

The observation period covers seven complete local days, September 14–20, 2026, in the Europe/Madrid time zone. Its exclusive endpoint is midnight at the start of September 21. The publication date is September 22; that date does not add another day of observations.

Across all included channels, the sample with citations contains 2,481 responses, 145 prompts, 24 projects and 21 accounts. The 558 OpenAI API responses form part of that collection. Differences between channels are descriptive and can reflect different prompts and collection methods. They are not an experiment using matched questions, a quality ranking or a measure of market share.

How the denominator was built

Collection included only completed executions and non-branded prompts, excluding accounts with “demo” in their email address. The latest response for each execution and stored model was retained, provided its text was nonempty and not an error. Citations came from citation records with a known collection method, rather than any web address detected in the text.

There were 3,026 valid responses from known channels. Citation telemetry was missing for 35, leaving 2,991 measured responses; 2,481 had valid citations. The headline denominator includes only responses containing at least one valid citation. Empty or missing citation records are excluded from it, so 18.46% does not describe the frequency across all generated responses.

What the finding supports

Published segments must contain at least 100 responses, ten prompts, three projects and three accounts; qualifying cases must span at least three accounts. These thresholds restrict publication of small groups, but they do not correct unequal weights or turn Bee LLM’s monitored activity into a sample representing all AI use.

The supported conclusion is specific: citing a single site was a recurring pattern in this OpenAI API collection, and its frequency changes when prompts receive equal weight. An absent or suppressed channel does not mean zero cases. No personal data or customer, project, prompt or response content is published, and the finding cannot support promises that publishing on a platform will secure citations.

Frequently asked questions

Does the 18.46% include responses without citations?

No. It is the percentage of the 558 sampled OpenAI API responses with valid citations that cited exactly one hostname.

Does a single site mean a single source document?

Not necessarily. Several pages and citations from the same site can meet the measure. The count does not establish how many independent documents support the response.

What does the 14.43% calculation add?

It shows the result when every prompt receives equal weight. Its difference from 18.46% reveals that prompt repetition influences the primary rate.