Skip to content

Knowledge Base

AI recommendations: 0.1 of a point decides whether your brand wins

Louie Valkhof
Louie Valkhof
16 min read
Isometric 3D illustration of a scale on which a brand logo balances against a single half star

A language model does not pick your brand because it is your brand. It picks your brand when there is nothing else to pick on. That is no longer an opinion, it has been measured: in an experiment with 670 comparisons where only the brand name differed the real brand was recommended every time, while nine controlled fake brands with exactly the same specifications were never named once. Under fair treatment every product should have scored around 10%.

That sounds like bad news for anyone who is not the market leader. The opposite is true. The moment there is something to choose on, the advantage of the familiar name almost entirely evaporates. In the same study the choice flips at a rating difference of 0.075 of a point, and brand identity explains no more than 1.2% of the variation in the ranking.

That shifts what positioning is about. Not "why we are different" in adjectives, but: which difference can a machine lay next to your competitor's and read off. The five studies below are all free and public, and they point the same way.

How much does your brand name weigh when an AI recommends a product?

Very little, as soon as there is something else to compare. Across 14,395 recommendations the researchers took apart where the variation in the ranking came from. The result is strict:

What determines the ranking Share of the variation
Product attributes (price, rating, review count, specifications) 82.4%
Position in the list as presented 6.5%
Interaction between factors 9.3%
Brand identity 1.2%

That 1.2% is not zero, and the nuance sits exactly there. When the information is identical, that 1.2% is all there is, and then the familiar name wins unconditionally. The authors therefore call brand preference a tiebreaker in so many words: it counts for most when the model has nothing else to steer by. And that happens more often than you would like, because most product pages in a category say roughly the same thing.

It is the same mechanism we already know from people. A buyer looking at two options that do not differ on paper falls back on recognition. The difference is that a model has no patience and no feeling: it takes the first hard difference it can read. We wrote earlier about who describes your brand when agents are reading along. This research adds the proportions. The description counts. The measurable facts around it count far harder.

Watch the scope. This is a laboratory setup with invented brands and specifications supplied inside the prompt, not a measurement of real webshops. It tells you how a model weighs things up, not how much revenue that moves.

Where exactly is the tipping point?

At the smallest difference the researchers could come up with. Across 9,220 comparisons they tested what it takes for an unknown brand to beat a real one. With completely identical product data the unknown brand won in 3.6 to 6.0% of cases. With the smallest advantage tested that jumped to 64 to 80%.

That smallest advantage is one of these three:

Signal Difference needed
Rating 0.075 of a point higher, less than a tenth
Review count 1.6 times as many
Price 7.3% lower

The authors sum it up themselves: less than the gap between two consecutive tenths of a star rating is enough to cancel out the entire brand advantage. That is a difference you cannot see in a screenshot and can barely see in a spreadsheet.

There is a counterpoint inside the same research that you have to read alongside it, otherwise the picture is wrong. In a different setup, where a fake brand was given better specifications across 2,769 attempts, the model still chose the real brand in roughly 96% of cases. Two seemingly opposite results, one explanation: the type of signal matters. A loose claim about specifications weighs light. A number that sits next to the competitor's as a comparable quantity, such as a rating or a price, weighs heavy. The model would rather compare figures than assertions.

A practical rule follows from that. Put your distinction into a quantity that fits next to your competitor's. Not "exceptionally sustainable", but the number on which you are more sustainable. Not "many satisfied customers", but the review count and the rating, visible on the page and in your structured data. That is the same argument we make about listings where the brand comes before the tactics, only now with a threshold value underneath it.

And there is a difference between models you need to know if you run tests: one model is far easier to flip than another. In this setup Claude was the most stubborn and GPT and Gemini were considerably more movable. Test on one model and you measure that model.

Why is writing for AI an entry ticket and not a lead?

Because the moment your competitor does the same, it is gone. That is not theory, it sits as a table in the same paper. The researchers let more challengers use authority language step by step and measured how often the market leader stayed standing, across 4,745 valid attempts:

Challengers using AI-targeted copy Market leader stays first choice
0 100.0%
1 19.8%
3 19.4%
6 73.1%
9 93.8%

Read that table top to bottom and you see a dip: it starts high, collapses, and then climbs back to just under where it began. One challenger getting its copy in order destroys the market leader's monopoly. As soon as everybody speaks the same language, that language loses its distinguishing power and the model falls back on what it has left: brand recognition. So the market leader does not win at the end because it got better, but because the distinction has been polished away.

Two things follow from this, and they do not contradict each other. First: not taking part is not an option. Across all 4,745 attempts the non-optimised brands received zero recommendations. Zero. Second: there is no strategy to build on it. The first mover earns a payoff of +0.802 in this measurement, and under general adoption that drops to +0.007. That is the definition of a catch-up race, not of a lead.

What does last is precisely what cannot be copied: real reviews, real specifications, real mentions at independent sources. The basic hygiene of your pages belongs in there too, and that sits in the seven elements of a category page an AI can read. But do not call it a strategy.

One more thing we should say out loud. The research also tests invented evidence claims as a tactic, and measures that they work. We do not do that, and that is not a tidy line for the bottom of a page. An invented claim that convinces a model is a claim a customer discovers later. We wrote earlier about why AI content costs you rankings when the substance is wrong; the same holds here, except the consequences are bigger because a recommendation is delivered with authority.

Which pages get cited, and why those?

The best structured page, not the best known name. That is a different level of the same game: first an engine has to use your page as a source, only then does what is on it count. A second study looked at exactly that. The authors collected the citations three engines gave on 70 product-focused questions and held the cited URLs against a framework of sixteen quality pillars. Every page got a single score out of that, between 0 and 1. The result per engine:

Engine Average quality score of the cited pages
Brave Summary 0.727
Google AI Overviews 0.687
Perplexity 0.300

The relationship is firm: page quality predicts citation with an odds ratio of 4.2. The pillars that stand out most are metadata and freshness, semantic HTML, and structured data. Those are not creative choices. That is construction work.

The researchers even hand you a working threshold: a score of 0.70 or higher combined with at least twelve pillar hits. That is more usable than the advice to "make your content AI friendly", because you can hold a page against it and see whether you sit below the line.

The difference between those engines is the practical part. Perplexity cites pages with a much lower average score, which means the bar is lower there. For a small Dutch brand that is the realistic first way in, while Brave and Google only join once the foundation is right. How you lay that foundation is in GEO for webshops.

Here too, the limitation belongs with it: seventy questions, one sector, English language, and the authors call their own setup observational. Take the mechanism with you, not the exact percentages.

Where do AI search engines get their evidence?

Not from you. AI search engines take their evidence from parties you do not control. A study out of Toronto set 1,000 product questions loose on AI search engines and on Google, and classified every cited source as brand owned, independent or social. For one category they publish both sides in full, cars in Canada:

Source type Google AI search engines
Independent (media, reviews, comparison sites) 40.6% 69.1%
Brand owned 36.6% 30.9%
Social (forums, video, community) 22.8% 0%

For software and consumer electronics they give only the AI side: 74.2 percent and 77.6 percent independent. For software we do know that Google pulls the brand forward instead, with 53.8 percent brand owned sources. So the order at AI search engines is independent above brand above social, and at Google it is exactly the other way round.

For small brands it is more extreme: there ChatGPT cited 95.1% independent sources in this measurement. Your own website is the foundation, then, but it is rarely the evidence.

There is a second finding in there that matters more for the Netherlands than the first. The researchers translated a hundred questions into five languages and saw that the overlap in cited domains between languages was low: the model switches to a different system of sources per language. For us that is the most important rule out of this research, and at the same time the most uncertain one. Dutch was not in the test. What you may assume is the mechanism: if the source system shifts per language, then Dutch language independent sources determine your visibility on Dutch language buying questions. What you may not assume is that this space is empty. You have to count that yourself per category, and that is exactly what we do before we spend a euro on it.

In practice that changes what you buy. Not a link building package, but a map of the Dutch domains that actually get cited in your category, plus the route to them. What that produces in measurable traffic is a separate story, and it sits in measuring AI referrals in GA4.

The citation rule belongs here too, and it is sharper than with the previous study: this measures the United States and Canada. That the source system shifts per language is yours to take. Which Dutch domains those are, nobody knows, because nobody has counted them.

Why does "the AI recommendation" not exist?

Because the models largely disagree with each other. A fourth measurement put 250 brand-free category questions to three models, five times per question, good for 3,750 answers across fifty brands in five sectors. The result that touches every visibility report: the models agreed on which brand was recommended best in only 41.6% of cases.

Put differently: in almost six out of ten categories a different brand sits at the top as soon as you look at another model. So a report selling you "your ChatGPT score" mostly measures ChatGPT.

Two other figures from the same measurement are good news for challengers. The concentration is moderate, not winner takes all: the measured inequality stays well under the threshold above which one party dominates the field. And in 8.0% of the questions there is no clear leader at all. Those are categories where nobody is the default choice and where you have something to claim.

For the way we work this has one hard consequence. A brand audit that measures AI visibility tests at least three models, with repeats per question, and reports the spread instead of one number. Do that differently and you deliver a snapshot of a single model, which the client will later call unreliable, rightly.

Why is an AI mention worth more than the click behind it?

You lose clicks. Roughly half of them. The only measurement in this file that rests on real human clicking behaviour rather than on model output comes from Pew, which followed the browsing behaviour of 900 Americans across 68,879 unique searches. Read it with the limitation attached: this is American, Google only, and the data is from April 2025.

Behaviour With AI summary Without
Click on a regular search result 8% 15%
Click on a source inside the summary 1% not applicable
Session ends afterwards 26% 16%

So the chance of a click roughly halves, and the source links inside the summary are barely used at all. That leads to the most important operational conclusion in this whole piece: the mention itself is the return, not the click. Adjust your reporting to that, otherwise you measure a decline while you are winning. What that means for your visitor numbers we wrote about earlier in agentic search and your webshop traffic.

One more figure that touches your content planning: the chance of an AI summary hangs strongly on the length of the question. At one or two words it happened in 8% of searches, at ten words or more in 53%. So short keywords barely pull that layer in; full sentence questions do. That is exactly the shape in which we build our knowledge base articles, and it is the reason we phrase headings as questions.

How do you test this yourself in an hour?

By repeating it on a small scale. You need no tool and no budget, just a short list of questions and discipline in writing things down. We do this at every brand scan and it almost always turns up something the client did not know.

Step one: write down five buying questions the way a customer asks them, in whole sentences. Not "best office chair" but "which office chair is best for someone who works from home eight hours a day and has back problems". That shape is not a style choice: the chance of an AI answer appearing rises along with the length of the question.

Step two: put every question to three models, and ask it five times. That last part feels excessive and is not. These models are not deterministic, so one answer is a sample of one. Note per run which brands are named and in what order.

Step three: note per answer which sources are cited. Count how many of those come from brands themselves and how many from independent parties. That ratio is your real work list, because it says where the recommendation gets built.

Step four: lay your product page next to those of the two brands that come up most often, and look for the first hard difference a machine can read. Rating, review count, price, a specification in structured data. If that difference is not there, you have your answer: there is nothing to choose on, and then the best known name wins.

What you have afterwards is not a score but a list. Five questions, three models, fifteen answers per question, and a source split per question. That is exactly the setup the studies above use at scale, and it takes an hour for one category. Most agencies would rather sell you a dashboard. This count is more accurate, because it covers your category in Dutch, and that is precisely the space the published measurements say nothing about.

What we do with this

Three things, and all three are duller than a tool with a score.

The first is the machine readable difference. In every brand project we look for the distinction that fits next to the competitor as a quantity, and we put it on the page and in the structured data. No adjectives, but a number, a certificate, a warranty term, a material specification. That is the translation of positioning into something a machine can compare, and it is the only intervention in this whole article for which we know a threshold value.

The second is the map of independent sources. Which Dutch domains get cited by the engines in your category, which of those are reachable, and in what order. That is work for our SEO service and it replaces the classic link building story.

The third is the measurement rule: three models, repeats, report the spread. No single number without the bandwidth around it. That is less sellable and far more usable.

What we do not do: invented claims, because they are demonstrably effective and that is exactly why they are a line. And we do not promise a position in an AI answer. The models agree on the top brand in only 41.6% of cases, and the source system shifts per language. So position differs per model and per language, and how stable it is over time nobody has measured. What we can do is lay the foundation that makes such a position possible in the first place. If you want to know where your brand stands now, get in touch and we will walk through your category.

One thing we cannot measure, and neither can anyone else. None of these five studies covered a Dutch category. We are running that count per client now, and once enough categories are in we will publish the measurement here. Until then the mechanism holds, not the percentage.

Louie Valkhof
Louie ValkhofFounder & Art Director, Oase Creative
Knowledge Base

Frequently asked questions

Need help with execution?

From strategy to production. Tell us about your project, no strings attached.

Response within 24 hours.
Start a project?