Skip to content
Original study

Which Sources AI Reuses When the Same Prompt Is Repeated

A frozen Bee LLM cohort shows that source reuse varies materially by engine, while repeated answers still draw from a changing pool of recorded hosts.

Bee LLM Observatory9 min read

AI engines do reuse part of their recorded source pool when the same prompt is run repeatedly, but the degree of reuse differs substantially by engine. In the eligible Bee LLM cohort, the prompt-equal mean share of unique recorded hosts that appeared in at least two responses was 51.8% for ChatGPT, 67.5% for Gemini, 57.9% for Perplexity, 75.9% for Claude and 39.5% for Google AI Mode.

Those percentages do not mean that the engines showed the same visible citations on every run. The measured object is a normalized host found in recorded source telemetry. It is evidence about recurrence in the collected source pool, not proof that every host was visibly cited, consulted in the same way or responsible for a statement in the answer.

The result at a glance

EngineEligible responsesEligible promptsMean share of unique recorded hosts reused
ChatGPT1,0146551.8%
Gemini5865167.5%
Perplexity5504557.9%
Claude1913175.9%
Google AI Mode1021639.5%

Each value is an average across prompts in which every eligible prompt carries equal weight. It is not a pooled percentage across all responses. That distinction matters because the cohort contains more ChatGPT observations than AI Mode observations and because some prompts were executed more often than others.

What the study asked

The practical question was narrow: when an identical prompt is repeated in the same AI engine, how much of the set of recorded source hosts returns in more than one answer? Monitoring teams often see a source once and treat it as a stable part of the engine's source landscape. The study tests whether that assumption is reasonable at host level.

It does not ask whether a domain is authoritative, whether a cited page caused the model to produce a claim, or whether one engine has a better retrieval system. The provider integrations expose different source objects, and the analysis does not equate a recorded URL with a visible attribution. Keeping the research question constrained prevents an operational telemetry study from becoming an unsupported ranking study.

The frozen cohort

The snapshot was frozen at 18:00 UTC on 21 July 2026. The broader non-demo production cohort contained 5,313 responses from 1,416 completed executions, 75 prompts, eight accounts and 11 projects. The recurrence analysis then applied stricter eligibility rules to each engine.

Demo accounts, incomplete executions, empty or errored responses, and records without at least one convertible HTTP or HTTPS source host were excluded. A prompt-engine cell needed at least five valid source-bearing responses. These rules produced 1,014 eligible ChatGPT responses across 65 prompts, 586 Gemini responses across 51 prompts, 550 Perplexity responses across 45 prompts, 191 Claude responses across 31 prompts and 102 AI Mode responses across 16 prompts.

Every public engine cut exceeded the minimum gates of ten executions, five prompts, three projects and three accounts. No customer, project, prompt, brand or host name appears in the aggregate artifact. Identifiers and text remained in remote memory during the read-only extraction.

How source reuse was defined

URLs were converted to lowercase hosts, the common www. prefix was removed, known internal wrappers were discarded, and a host was counted at most once inside one response. For a prompt and engine, the analysis first built the union of all unique recorded hosts. It then marked a host as reused when it appeared in at least two eligible responses.

The main metric divides reused unique hosts by all unique hosts observed for that prompt-engine cell. Imagine five eligible answers containing 20 distinct hosts in total. If eight of those hosts occur in two or more answers, the reuse share is 8 divided by 20, or 40%. It does not matter whether a repeated host occurs twice or in all five answers for this particular metric.

Two sensitivity measures help interpret the headline. The repeat-incidence share asks what portion of all response-host incidences came from reused hosts, so frequently recurring sources receive more influence. Mean pairwise Jaccard similarity compares the host set in every pair of responses. Both measures were retained in the frozen aggregate, but the unique-host reuse share remains the clearest answer to the stated question.

Why prompt-equal weighting matters

A raw pooled calculation would let high-frequency prompts dominate. If one prompt had 100 responses and another had five, the first could determine most of the engine result even if both represented equally important customer questions. The study calculates metrics within each eligible prompt first and then averages those prompt results.

This choice changes what the number means. The 51.8% ChatGPT result describes the typical pattern across the 65 eligible prompts under equal prompt weighting. It is not the percentage of all ChatGPT host rows that repeated. For a content or measurement team, prompt-equal weighting is usually more useful because it keeps the designed prompt inventory, rather than execution volume, as the analytical frame.

How to read the differences between engines

Claude recorded the highest mean reuse share in this cohort at 75.9%, while AI Mode recorded the lowest at 39.5%. Gemini and Perplexity sat between those values, and ChatGPT was close to half. These are descriptive differences under the observed integrations and eligibility rules. They are not a league table of engine quality.

The size of the eligible pools also differs. Claude's result is based on 31 prompts and AI Mode's on 16, compared with 65 for ChatGPT. Source telemetry has different semantics across providers. A host captured by one integration may represent a different stage or display condition than a host captured by another. The safe interpretation is that the operational recurrence pattern varied by engine in this cohort, not that one engine consistently relies on fewer or better sources.

Variation also exists within an engine. ChatGPT prompt-level reuse shares ranged from 22.9% to 71.9%, with a median of 51.4%. Gemini ranged from 36.0% to 95.0%, and Perplexity from 33.3% to 83.3%. A single engine average therefore cannot replace a prompt-level diagnostic. Category, phrasing, time and available search results can all accompany different source pools.

What the results mean for visibility measurement

A one-run source list is incomplete evidence. Even in the engine with the highest aggregate recurrence, not every unique host reappeared. Teams that collect only one answer can mistake a transient source for a stable one or miss a source that returns in later executions. This is the same sampling problem that affects mentions, recommendation positions and answer wording.

Repeated collection should preserve atomic observations before aggregation: exact prompt, engine, locale, run time, valid or failed state, recorded hosts and visible citations where the integration exposes them. A monitoring system should not merge detected URLs and explicit citations into one field. The distinction makes it possible to analyse source discovery without claiming visible attribution.

For a broader measurement design, use the framework for AI search visibility metrics and KPIs. If the decision concerns brand presence rather than sources, the workflow for tracking brand mentions in AI search explains aliases, prompt sampling and repeated observations.

A practical recurrence workflow

  1. Freeze the prompt inventory. Separate category, comparison and problem prompts so a changing mix does not masquerade as changing source behaviour.
  2. Repeat under documented conditions. Record engine, locale, mode, location and collection time. Do not combine incompatible product surfaces without a label.
  3. Keep missing states explicit. A failed response or absent telemetry is not an empty source set and must not become zero recurrence.
  4. Normalize conservatively. Deduplicate hosts within a response and document prefix or wrapper removal. Do not collapse unrelated subdomains without a business rule.
  5. Calculate within prompt first. Inspect recurrence, incidence and pairwise overlap per prompt before producing an engine summary.
  6. Review source examples privately. Aggregates identify where recurrence is unusual; authorized analysts can then inspect source-level examples under the appropriate privacy controls.

This workflow makes changes auditable. If recurrence falls after an engine or integration update, the team can determine whether the shift came from answer behaviour, the prompt mix, a collection failure or a telemetry definition.

What this study cannot establish

The sample is not a census of all prompts, industries or users. It contains the non-demo Bee LLM projects that met the eligibility rules by the cutoff. Equal prompt weighting reduces execution-volume bias within the cohort, but it does not turn the cohort into a representative market panel.

The analysis operates at host level. Two different pages on the same host are treated as the same source, while closely related subdomains can remain separate. It cannot tell whether a page's content changed or whether the same passage supported an answer. It also does not measure source quality, audience reach, referral traffic or causal influence on a recommendation.

Most importantly, recorded source telemetry is not automatically visible citation telemetry. The phrase “reused source” throughout this article means a reused normalized host in the eligible recorded-source field. Claims about citation visibility require an explicit citation map and a separate denominator.

Reproducibility and integrity

The extraction ran against Bee LLM's production PostgreSQL data in a repeatable-read, read-only transaction. The frozen aggregate, methodology, tabular export and analysis wrapper are hashed. The aggregate linked to this article has SHA-256 dac3a0a862b513789744b34e2879c4bc847fcc3dfd989505f40420aa0468e354. Recomputing or editing the file would change that digest and fail the editorial evidence gate.

The main value of the result is not a universal percentage. It is the demonstrated need to treat sources as repeated observations. In this cohort, every engine reused part of its recorded host pool, and every engine also introduced hosts that did not persist across all runs. A defensible visibility programme should measure both recurrence and change instead of treating one generated answer as permanent evidence.

Frequently asked questions

Does a recorded source host mean the user saw a citation?

Not necessarily. This study uses hosts recorded in detected source telemetry. A recorded host can differ from a citation visibly attributed in the answer, so the results describe source recurrence rather than visible citation recurrence.

Why did every prompt receive equal weight?

Some prompts had many more eligible responses than others. Giving each prompt equal weight prevents a small number of frequently executed prompts from dominating an engine-level result.

Can these percentages predict what a specific brand will see?

No. They describe the eligible Bee LLM cohort through 21 July 2026. A particular brand, market, prompt or future engine version can show a different pattern, so teams should measure their own repeated sample.