Branded vs unbranded prompts: what your AI score measures
by John LeeBuilding peekr in Seoul, measuring how AI search engines name brands. Previously co-founded vlogr and shipped iOS apps (2018–2021).
Put a brand's name inside the question and the answer comes back with the brand's name in it. That is most of what a high AI visibility score can be made of, and it is easy to check.
We ran the comparison inside single brands, so nothing else moves. Four brands in our archive were named in 100% of the answers to questions that contained their own name. The same four brands, asked the category question a buyer would type, were named 26.9%, 37.1%, 6.3% and 4.2% of the time.
Same brand, same two engines, same measurement run. The only thing that changed was whether the question said the name.
Branded and unbranded prompts, defined
A branded prompt is a question that contains the brand's own name. "What does Acme do?" An unbranded prompt is the category question someone asks before they know the brand exists. "What's the best tool for editing podcast clips?" Both are legitimate to track. They measure different things, and only one of them is a test the brand can fail.
What we measured
Each figure below is one brand and one question type, pooled across both engines. A rate is the share of parsed answers in that group that named the brand. Denominators are in brackets, and they are small.
| Brand ◐ | Named — question contained its name | Named — category question |
|---|---|---|
| Brand 1 | 100.0% (12) | 26.9% (108) |
| Brand 2 | 100.0% (12) | 37.1% (89) |
| Brand 3 | 100.0% (12) | 6.3% (96) |
| Brand 4 | 100.0% (12) | 4.2% (72) |
| peekr (us) | 32.7% (55) | 0.0% (161) |
For brands 3 and 4 that is roughly a 16-fold and a 24-fold gap between the two columns. Every one of the four looks flawless when you hand the model its own name.
Why we are not naming the four brands
The four numbers in the right-hand column are not comparable with each other, and putting names on the rows would invite exactly the reading the data does not support.
Each brand in our archive carries its own list of questions — one has 26 of them, another 25, ours has 10 — and they are questions about different categories. A brand scoring 37.1% and a brand scoring 6.3% sat different exams. Ranking them would be the same error as averaging two engines into one number.
What survives is the comparison inside each row. Both cells in a row came from the same brand's question list, on the same day, through the same pipeline. That holds no matter how the list was written, which is why this post is five separate findings rather than one league table.
Why the gap exists
Mention detection is string matching: the name is in the answer text or it is not.
A branded question hands the model the string before it starts writing. Answering "what does Acme do" without the word Acme in the reply takes effort. An unbranded question asks the model to produce a name it was not given, from whatever it retrieved or stored. One question supplies the answer; the other requires one.
That is the mechanism as we understand it. We have not tested it — we measured the gap, not its cause.
The fifth row is us, and it breaks the pattern
Our own branded rate is 32.7%, not 100%. Asked a question with the word peekr in it, the engines named us in about a third of the answers. On the category question we were named in 0 of 161.
So "the branded question always returns you" is not a law. It is what happened for four brands and not for us. We do not know what separates them and we did not measure it. The most we will say is that a branded question is a weak test in both directions. It flatters brands the engines have something to say about, and in a separate measurement it followed the phrasing of the question more than it followed the company — a name we invented passed a leading self-check about as convincingly as the real companies did, in Does ChatGPT Know My Brand?.
What this does to a visibility score
A visibility score is a rate, and a rate is a weighted average over whatever questions are in the set. If the set leans branded, the score inherits the branded column. Two tools can measure the same brand on the same day and report numbers far apart, both of them arithmetically correct.
This is not an accusation. We have not audited anyone's question mix and cannot see one from outside. What we can say about the market is narrower and already published: of the eight vendor pages we read for 8 AI visibility tools compared, none documents how many times it asks each question.
One thing that follows, and it is a warning rather than a tactic: rewriting your question set moves the number without moving anything in the world. Swapping unbranded questions for branded ones is changing the ruler. If you do it, the series before and after are two different measurements and should not be drawn as one line.
What we report instead
We keep the two rates apart and print the denominator next to each. The split by question type is the first thing on the screen rather than a filter you have to find, because a single blended percentage is the one number in this category we know to be unreadable. The six-step version of that method, worked on a 336-answer dataset, is in How to track your brand in AI answers.
If your unbranded rate turns out to be zero, the next question is which stage failed — rank, retrieval or selection. That is a different investigation: Why your brand isn't in AI Overviews.
Run the free check on one URL — it reads the page, then asks one unbranded question about your category and shows the raw answer with its sources. One question asked once is a demonstration, not a baseline.
Four things to ask about any visibility number
- What share of the questions behind this number contain my brand name? If the answer is not on the screen, the number is not readable yet.
- Show me the two rates separately, each with its denominator. Ours are in the table above, including the row where we score zero.
- How many samples per question? In the same archive, whether a brand is named flips 11.8% of the time across 314 cells when nothing changes but the sampling ◐. A rate built on one draw per question is mostly noise.
- Which rows were eligible to be counted at all? That is the question that moved one of our own numbers from 12.7% to 37.7% — Check the denominator before you trust an AI metric.
Frequently asked questions
What is a branded prompt in AI search tracking?
A question that contains the brand's own name, such as "what does Acme do" or "is Acme any good". An unbranded prompt is the category question a buyer asks before knowing the brand exists. In our measurement four brands were named in 100% of answers to branded questions and in 4.2% to 37.1% of answers to their category questions ◐.
Why is my AI visibility score high when no customer finds me through AI?
Check the question mix before anything else. A set weighted toward questions containing your name reports a high rate regardless of what the category answers look like — in our data the two rates differed by roughly 16 and 24 times for two of the four brands ◐. Ask for the branded and unbranded rates separately, each with its denominator.
Should I stop tracking branded prompts?
No. They answer a real question: what an engine says about you when someone already has your name and is checking you out. Errors in that answer are worth knowing about. The mistake is merging the two rates into one headline number, because the branded half is close to a constant for brands the engines can retrieve.
Does this table show which brand is most visible in AI search?
It does not, and we would resist reading it that way. Each brand runs its own question list about its own category, so the right-hand column is five separate exams rather than one scoreboard. The finding is the distance between the two columns within each row.
How many samples per question do I need?
More than one, and the tool should tell you how many it uses. In the same 1,075-answer archive, whether a brand is named flips 11.8% of the time across 314 cells with the question, engine and run held fixed ◐. That is the floor of noise underneath any single reading.
Method. Our own stored measurements, read 28 August 2026: 11 brands and 1,075 parsed answers in the archive, two engines (ChatGPT and Gemini), three samples per prompt, raw responses retained. The within-brand comparison could be run on five of those brands. Each rate is the share of parsed answers in that brand-and-question-type group that named the brand, pooled across both engines, and every denominator is printed in the table. The 11.8% mention-flip figure is measured over cells — one run × prompt × engine with two or more samples — where brand_mentioned was not constant, across 314 cells.
Limits. Small, and small in a way that matters: 12 branded answers per brand for four of the five rows. One vendor's pipeline, our own parser, no control group. Rates are pooled across two engines and we did not split this comparison by engine, although our own earlier data says the engines do not run the same mechanism. The brands were not selected at random — they are the brands already in our archive. Nothing here is causal: we observed that question type and naming rate move together, and we did not test why. Two of the four unnamed brands sell in our own category, and the fifth row is us — we have an interest in how this reads.
Source. Every figure in this post is our own first-party measurement ◐. No external dataset is used. The sentiment half of the same archive, including a mistake we made counting it, is in We removed sentiment from our dashboard.