Skip to content
Original study

How Often ChatGPT and Gemini Cite the Same Domains

Across matched answers with explicit citation telemetry in both engines, most pairs shared no citation host; source definitions materially changed the result.

9 min read

ChatGPT and Gemini often recorded different explicit citation hosts for the same prompt in the same execution. Among 337 eligible matched executions, 204 pairs — 60.5% — shared no explicit citation host. After calculating the rate within each prompt and giving all 71 prompts equal weight, the mean zero-overlap rate was 61.4%.

The result is conditional. Both responses had to contain a non-empty explicit citation map, so it does not represent runs without that telemetry. It also changes materially when the source definition changes: using the broader detected-URL field, 31.2% of pairs had no host overlap. Explicit citations and broadly detected sources are related but different analytical objects.

The paired result

MeasureExplicit citation hostsBroad detected URLs
Matched executions337337
Pairs with zero host overlap204105
Pair-level zero-overlap rate60.5%31.2%
Median pairwise Jaccard00.0357

The explicit-citation Jaccard interquartile range was 0 to 0.0833. A Jaccard value of zero means the two normalized host sets had no element in common. A value of one would mean identical sets. The low median therefore describes limited host overlap, not whether either answer was accurate or useful.

What the study compared

The study matched a valid ChatGPT response and a valid Gemini response produced for the same execution and prompt. Matching within an execution holds the stored prompt and execution context together more tightly than comparing unrelated answers from different days. The analysis extracted only hosts linked to answer fragments in each engine's explicit cite_map, normalized them and removed duplicates inside each response.

For every pair, it calculated the intersection and union of the two host sets. The Jaccard score is the size of the intersection divided by the size of the union. The analysis also recorded whether the intersection was empty. Because some prompts had more eligible matched executions, the headline rate was calculated within prompt first and averaged across prompts.

This design answers a precise question: under the eligible telemetry conditions, how often did ChatGPT and Gemini expose at least one of the same citation hosts for matched prompts? It does not attempt to determine which system found a source first, whether a host influenced generation, or whether the citation appeared with equal prominence.

Eligibility and the frozen window

The evidence was frozen at 18:00 UTC on 21 July 2026. Eligible matched pairs span 13 July at 07:45:45 UTC through 21 July at 15:50:03 UTC. The final set includes 337 executions covering 71 prompts, 11 projects and eight accounts.

Demo accounts, incomplete executions, empty or errored responses, and any pair in which one engine lacked a non-empty explicit citation map were excluded. Absence of a citation map is missing telemetry for this analysis, not evidence of zero citations. Converting those missing cases to empty sets would inflate the zero-overlap rate and answer a different question.

The cohort passed the predefined public thresholds for accounts, projects, prompts and executions. The extraction ran in a repeatable-read, read-only database transaction. No prompt text, host, brand, project or account identifier was emitted into the frozen aggregate.

Why the prompt-equal result is primary

The raw 60.5% figure treats all 337 matched executions equally. That is a valid pair-level description, but prompts with more repeated runs receive more influence. The 61.4% headline instead calculates the zero-overlap share separately for each prompt, then averages the 71 prompt shares.

The two results are close in this cohort, which is reassuring, but the prompt-equal method remains the defensible primary measure. A monitoring prompt inventory represents a set of questions, not a lottery in which the most frequently executed question should define the whole programme. Prompt-level clustering also means that 337 pairs should not be interpreted as 337 fully independent topics.

Across prompts, the zero-overlap rate ranged from 0% to 100%, with a median of 66.7% and an interquartile range from 33.3% to 100%. That spread is operationally important. Some prompts produced more shared hosts, while others repeatedly separated the engines' explicit citation pools.

Why source definition changes the answer

The broad sensitivity analysis used detected_urls rather than hosts explicitly linked to answer fragments. In that broader object, 105 of 337 pairs had no overlap, or 31.2%, and the median Jaccard score rose from zero to 0.0357. The difference is too large to hide in a footnote.

A detected URL can represent source discovery or provider telemetry that is not displayed as an explicit answer attribution. An explicit citation map connects a host to a fragment in the recorded response structure. Neither object is automatically “better”; they answer different questions. If a dashboard merges them, it can make citation overlap look much higher without any actual change in visible attribution.

Measurement teams should name the field in every chart. “Citation-host overlap” should be reserved for explicit attribution telemetry. “Detected-source overlap” can describe the broader set. If an engine or integration stops providing one field, the series should be marked unavailable rather than silently replaced with the other.

What low overlap does and does not imply

Limited overlap means that cross-engine source visibility is not interchangeable in this sample. A domain observed in ChatGPT cannot be assumed to appear in Gemini for the same question, even when both answers expose explicit citation telemetry. A source-monitoring programme that samples only one engine can therefore miss a substantial part of the citation landscape seen in another.

It does not mean one answer was wrong. Two engines can support compatible conclusions with different sources. Conversely, sharing a host does not prove identical reasoning, equal prominence, correct interpretation or a good user experience. The metric says nothing about source authority, page freshness, sentiment, answer completeness or conversion value.

It also does not prove persistent engine preferences. The window covers just over eight days and the results describe the eligible Bee LLM cohort. Search results, model behaviour, provider interfaces and the underlying web can change. A longitudinal claim requires repeated frozen panels under comparable conditions.

Implications for brand and content teams

First, measure engines separately. A single combined “AI citations” count can conceal that different surfaces expose different domains for the same prompt. Preserve engine, prompt, locale, execution time and source type at observation level before building totals.

Second, distinguish presence from attribution. A brand can be mentioned without an owned-domain citation, and a page can be cited without the brand becoming a recommendation. The broader framework for AI search visibility metrics and KPIs keeps these dimensions separate.

Third, repeat the sample. This study compares engines within matched executions, but one matched pair remains one observation. Repeating the same prompt reveals whether a citation host is recurrent or transient. The guide to tracking brand mentions in AI search covers the same sampling discipline for brand presence.

Finally, investigate gaps rather than chasing every domain. If an important prompt consistently exposes category references in one engine but not another, inspect whether the missing source type reflects documentation, independent evidence, entity consistency or simply different telemetry. The observed difference should start a diagnosis, not produce an automatic content prescription.

A reproducible overlap workflow

  1. Define a matched key. Use the same prompt and execution context; reject ambiguous joins.
  2. Separate source objects. Store explicit citations and broadly detected URLs in distinct fields.
  3. Normalize at host level. Lowercase, remove a documented common prefix and deduplicate within each response.
  4. Keep missing telemetry missing. Do not turn an absent map into an empty set.
  5. Calculate pair metrics. Record intersection, union, Jaccard and zero-overlap state.
  6. Aggregate within prompt. Let each designed question contribute equal weight to the primary summary.
  7. Publish denominators and sensitivity. Readers need to see how a broader source definition changes the result.

This sequence makes the calculation auditable and prevents provider volume from deciding the outcome. It also allows a future integration change to be isolated: if explicit citation mapping disappears, the metric becomes unavailable instead of breaking continuity invisibly.

Limitations

The analysis is conditional on non-empty explicit citation maps in both engines. It cannot estimate citation overlap for excluded pairs. The engines may expose attribution differently, and a structurally linked fragment does not guarantee that every interface displayed the citation in the same way to every user.

Hosts are a coarse unit. Multiple pages from one domain collapse to one host, while related subdomains can remain separate. The study does not assess page-level overlap, quote-level support, source quality or causal contribution. It also does not publish domain names, so it cannot identify which categories of source explain the difference.

The sample is not a market-representative panel. It is a privacy-preserving aggregate of eligible Bee LLM production observations through a fixed cutoff. The prompt-equal design controls one form of weighting bias but cannot remove selection effects in the underlying projects and prompts.

Evidence integrity

The frozen aggregate associated with this article has SHA-256 5805e41a5dfc0158a67bb82b86128c7d3d469b0a1e834e94599aa1b6a7511cb2. The methodology, CSV export, analysis wrapper and aggregate are stored with checksums, and the extraction can be rerun through the common read-only evidence script.

The defensible conclusion is specific: in 337 matched executions with explicit citation telemetry in both engines, ChatGPT and Gemini frequently exposed no shared citation host, and the prompt-equal zero-overlap rate was 61.4%. The much lower 31.2% zero-overlap rate for detected URLs demonstrates why every citation analysis must state exactly which source object it measures.

Frequently asked questions

Did 61.4% of all ChatGPT and Gemini answers share no citations?

No. The 61.4% prompt-equal mean applies only to matched executions where both engines recorded a non-empty explicit citation map. It must not be generalized to answers without that telemetry or to all users.

Why is the detected-URL result different from explicit citation overlap?

Detected URLs form a broader source object than hosts explicitly linked to answer fragments. In the sensitivity set, 31.2% of pairs had zero overlap, compared with 60.5% at pair level for explicit citation hosts.

Does low overlap mean one engine used poor sources?

No. Overlap measures whether normalized hosts are shared, not source quality, correctness or causal influence. Two accurate answers can cite different domains, while shared domains do not guarantee equal answer quality.