Skip to content

Half of small brands never appear once in 37,000 AI answers

Researchers audited 37,000 AI assistant answers recommending suppliers across 19 sectors, using a catalogue of 533 brands. Between 48% and 52% of the small and regional brands appeared in none of them. Not once.

·7 min read·Research

When somebody tells you their brand "does not come up much" in ChatGPT, you assume they mean rarely. This study says something else: for half of small brands, the number is not low. It is zero.

And it brings a second surprise, this one for the big players: being the leader in your category does not get you recommended either.

What they did

The team ran about 37,000 queries against AI assistants using 215 prompts of the kind a real buyer writes: which supplier to use, which product to pick, which company to hire. Nineteen different sectors.

The clever part came first. They built a catalogue of 533 brands and sorted them into five levels by how well known they are, from global category leader to regional specialist. That let them go past "the AI names some and not others" and see the pattern band by band.

37,000AI answers analysed, using 215 commercial recommendation prompts
533Brands in the catalogue, split into five levels by size
19Different sectors, so the pattern does not hang on one market

The hard finding: invisibility is not low visibility

In the two lowest levels, the specialists and the regional players, between 48% and 52% of brands did not appear in any of the 37,000 runs. The authors call it catastrophic invisibility and the name fits.

It is worth pausing on what that means. It is not that they come up rarely and need improving. It is that across 37,000 chances, 215 different questions and 19 sectors, the model never mentioned them at all. For practical purposes they do not exist in that channel.

Small and regional brands that never appeared once

Across the study's 37,000 runs

Never appear48 to 52 out of 100
Appear at least once48 to 52 out of 100

Source: W. Jack, N. Lehman, K. Maloney and S. Xu, Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit, 2026.

And now the surprise for the big players

You would expect the category leader, the one that shows up everywhere, to take nearly all the recommendations. It does not.

Top-level brands appear in almost every option set the model gets to work with, so finding them is not the problem. But of the slots they reach, they only take between 25% and 41%. Most of the time they are available, the model ends up recommending somebody else.

The level that converts best is not the first, it is the second: strong brands that are not the number one win between 37% and 52% of the times they get there, the highest rate in the whole study. They are known enough to be in the conversation and specific enough to fit a concrete need.

Out of every 100 times it reaches the recommendation, how many it wins

By brand size level

Strong brands that are not the number one37 to 52
Mid-sized brands34 to 40
Category leaders25 to 41

Source: the same audit. Bars are drawn on the midpoint of each range. For mid-sized brands coverage already drops to 88%, meaning they start not being available at all.

The reading is uncomfortable for everyone, in different ways. If you are small, your problem is that they cannot find you. If you are large, they find you and still pick somebody else. Those are two different jobs.

The three reasons you get left out

The same authors published a second study that completes the first and is, in practice, more useful. Using 215 commercial recommendation prompts they compared ChatGPT and Claude and sorted every failure into one of three moments.

1. It cannot find you

Your brand never even reaches the set of options the model considers. This is the small brand failure and it explains total invisibility.

2. It finds you but you do not convince it

You are available, but when it compares, another option reads as more compelling. This is where leaders who are present still lose.

3. You convince it, but for a different customer

You fit, but not the specific profile in the question. This effect spikes when the person asking describes their situation in detail.

Here is the valuable part. When ChatGPT and Claude both leave a brand out, they agree on the reason 95.1% of the time, across 7,763 joint failures. And that agreement rises to 99.6% precisely for small and regional brands.

Translated into something you can act on: for a small brand, the diagnosis is reliable. If two different models leave you out for the same reason almost every time, you know what to fix without guessing. For category leaders that agreement drops to 81%, so there it really is worth looking engine by engine.

Out of every 100 times ChatGPT and Claude both leave you out, how often for the same reason

Across 7,763 cases where both failed to recommend

Small and regional brands99.6
All brands95.1
Category leaders81

Source: W. Jack, N. Lehman, K. Maloney and S. Xu, Divergent Recommendations, Convergent Diagnoses, 2026.

What to do, depending on where you are

The study offers no recipes, but it does make clear the recipe cannot be the same for everyone, because the failure is not the same.

The limits, which also need saying

Before drawing conclusions

  • It is a snapshot, not an experiment. It measures how things stand, it does not prove that changing X improves anything. Nobody has yet moved a brand between levels and measured the effect.
  • The levels are defined by the authors. Whether a brand is level 4 rather than level 3 is their classification, reasonable but arguable.
  • It is 215 prompts. Many runs, yes, but over one specific set of questions. Different questions would move the percentages.
  • It does not cover every engine. It compares ChatGPT and Claude style assistants. It says nothing about Google's AI Overviews or Perplexity, which behave differently.

Even with all that, the order of magnitude is hard to argue with. Half the brands in a 533-brand catalogue never appearing across 37,000 attempts is not a methodological quibble.

Questions people ask us

Does this mean there is nothing to do if I am small?

It means the opposite, but in a different order. The first goal is not getting recommended, it is entering the set of options the model works with. You do that by appearing on the pages it reads, which are almost never your own.

Why do leaders only win one in three?

Because being available and being the best answer to a specific question are different things. When somebody describes in detail what they need, the model looks for fit, not fame. That is where a more specific brand overtakes a bigger one.

How do I know which of the three failures I am in?

By reading the full answers, not just whether you appear. If the model does not name you and does not seem to know you, it is the first. If it mentions you in passing and recommends another, the second. If it describes you well but for a type of customer that is not yours, the third.

Does this hold outside the US?

With caution. The study does not publish a country breakdown and the brand catalogue is mainly international. The mechanism, prior fame deciding who makes the list, has nothing market-specific about it, but the exact percentages could move.

Find out which of the three you are in

Bee LLM asks your questions to nine AI assistants, stores the full answers and shows you whether they name you, who they name instead and which pages they build the answer from. Free plan, no card needed.

Start tracking free

Sources: W. Jack, N. Lehman, K. Maloney and S. Xu, Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit and Divergent Recommendations, Convergent Diagnoses, 2026. More studies like these, summarised daily, at GEO al día. Read next: how we measure whether an AI names your brand and why Google's top 10 is not enough.