Half of small brands never appear once in 37,000 AI answers
Researchers audited 37,000 AI assistant answers recommending suppliers across 19 sectors, using a catalogue of 533 brands. Between 48% and 52% of the small and regional brands appeared in none of them. Not once.
When somebody tells you their brand "does not come up much" in ChatGPT, you assume they mean rarely. This study says something else: for half of small brands, the number is not low. It is zero.
And it brings a second surprise, this one for the big players: being the leader in your category does not get you recommended either.
What they did
The team ran about 37,000 queries against AI assistants using 215 prompts of the kind a real buyer writes: which supplier to use, which product to pick, which company to hire. Nineteen different sectors.
The clever part came first. They built a catalogue of 533 brands and sorted them into five levels by how well known they are, from global category leader to regional specialist. That let them go past "the AI names some and not others" and see the pattern band by band.
The hard finding: invisibility is not low visibility
In the two lowest levels, the specialists and the regional players, between 48% and 52% of brands did not appear in any of the 37,000 runs. The authors call it catastrophic invisibility and the name fits.
It is worth pausing on what that means. It is not that they come up rarely and need improving. It is that across 37,000 chances, 215 different questions and 19 sectors, the model never mentioned them at all. For practical purposes they do not exist in that channel.
Small and regional brands that never appeared once
Across the study's 37,000 runs
Source: W. Jack, N. Lehman, K. Maloney and S. Xu, Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit, 2026.
And now the surprise for the big players
You would expect the category leader, the one that shows up everywhere, to take nearly all the recommendations. It does not.
Top-level brands appear in almost every option set the model gets to work with, so finding them is not the problem. But of the slots they reach, they only take between 25% and 41%. Most of the time they are available, the model ends up recommending somebody else.
The level that converts best is not the first, it is the second: strong brands that are not the number one win between 37% and 52% of the times they get there, the highest rate in the whole study. They are known enough to be in the conversation and specific enough to fit a concrete need.
Out of every 100 times it reaches the recommendation, how many it wins
By brand size level
Source: the same audit. Bars are drawn on the midpoint of each range. For mid-sized brands coverage already drops to 88%, meaning they start not being available at all.
The reading is uncomfortable for everyone, in different ways. If you are small, your problem is that they cannot find you. If you are large, they find you and still pick somebody else. Those are two different jobs.
The three reasons you get left out
The same authors published a second study that completes the first and is, in practice, more useful. Using 215 commercial recommendation prompts they compared ChatGPT and Claude and sorted every failure into one of three moments.
Your brand never even reaches the set of options the model considers. This is the small brand failure and it explains total invisibility.
You are available, but when it compares, another option reads as more compelling. This is where leaders who are present still lose.
You fit, but not the specific profile in the question. This effect spikes when the person asking describes their situation in detail.
Here is the valuable part. When ChatGPT and Claude both leave a brand out, they agree on the reason 95.1% of the time, across 7,763 joint failures. And that agreement rises to 99.6% precisely for small and regional brands.
Translated into something you can act on: for a small brand, the diagnosis is reliable. If two different models leave you out for the same reason almost every time, you know what to fix without guessing. For category leaders that agreement drops to 81%, so there it really is worth looking engine by engine.
Out of every 100 times ChatGPT and Claude both leave you out, how often for the same reason
Across 7,763 cases where both failed to recommend
Source: W. Jack, N. Lehman, K. Maloney and S. Xu, Divergent Recommendations, Convergent Diagnoses, 2026.
What to do, depending on where you are
The study offers no recipes, but it does make clear the recipe cannot be the same for everyone, because the failure is not the same.
- If you never appear, your job is not to persuade, it is to exist. Be on the places the model draws its options from: sector comparisons, directories, reviews, forums. Writing your website better does not fix not being found.
- If you appear and do not get picked, you already have the hard part. What is missing is a concrete reason to prefer you, and the margin needed is surprisingly small: another study measured that less than a tenth of a star more in rating is enough to flip a comparison between equals.
- If you get picked for the wrong customer, it is a fit problem. Say who you are for and who you are not for, in those words. AI answers are built on descriptions, and a vague description puts you in the wrong conversation.
- And whichever it is, measure before acting. All three failures look identical from outside: you are not there. Only by reading the full answers and the sources behind them do you learn which of the three you are in.
The limits, which also need saying
Before drawing conclusions
- It is a snapshot, not an experiment. It measures how things stand, it does not prove that changing X improves anything. Nobody has yet moved a brand between levels and measured the effect.
- The levels are defined by the authors. Whether a brand is level 4 rather than level 3 is their classification, reasonable but arguable.
- It is 215 prompts. Many runs, yes, but over one specific set of questions. Different questions would move the percentages.
- It does not cover every engine. It compares ChatGPT and Claude style assistants. It says nothing about Google's AI Overviews or Perplexity, which behave differently.
Even with all that, the order of magnitude is hard to argue with. Half the brands in a 533-brand catalogue never appearing across 37,000 attempts is not a methodological quibble.
Questions people ask us
Does this mean there is nothing to do if I am small?
It means the opposite, but in a different order. The first goal is not getting recommended, it is entering the set of options the model works with. You do that by appearing on the pages it reads, which are almost never your own.
Why do leaders only win one in three?
Because being available and being the best answer to a specific question are different things. When somebody describes in detail what they need, the model looks for fit, not fame. That is where a more specific brand overtakes a bigger one.
How do I know which of the three failures I am in?
By reading the full answers, not just whether you appear. If the model does not name you and does not seem to know you, it is the first. If it mentions you in passing and recommends another, the second. If it describes you well but for a type of customer that is not yours, the third.
Does this hold outside the US?
With caution. The study does not publish a country breakdown and the brand catalogue is mainly international. The mechanism, prior fame deciding who makes the list, has nothing market-specific about it, but the exact percentages could move.
Find out which of the three you are in
Bee LLM asks your questions to nine AI assistants, stores the full answers and shows you whether they name you, who they name instead and which pages they build the answer from. Free plan, no card needed.
Start tracking freeSources: W. Jack, N. Lehman, K. Maloney and S. Xu, Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit and Divergent Recommendations, Convergent Diagnoses, 2026. More studies like these, summarised daily, at GEO al día. Read next: how we measure whether an AI names your brand and why Google's top 10 is not enough.
