Does ChatGPT Know My Brand? A Fake Name Scored 90.5%
by John LeeBuilding peekr in Seoul, measuring how AI search engines name brands. Previously co-founded vlogr and shipped iOS apps (2018–2021).
Do not audit your brand by asking ChatGPT about your brand. The answer follows the wording of your question more than it follows your company.
We measured it. Asked in a leading way, a company name we invented scored 90.5%. Asked skeptically, Beaconly — a real product shipping on the Mac App Store — scored 7.1%, barely above the invented name at 3.5%. Our own name moved from 97.7% to 12.9% with nothing changing except the phrasing.
The audit that does work is further down, as a table you can paste into a spreadsheet. It takes about an hour and needs no tool.
Key takeaways
- A leading prompt agrees with everything. Asked "Is X a tool for tracking brand visibility in AI search answers?", the lowest score any name received was 90.5% — and it belonged to a company we invented ◐.
- A skeptical prompt denies real companies. Beaconly, a product shipping on the Mac App Store, scored 7.1%; our invented name scored 3.5%. Those 3.6 points are the whole distance between "exists and ships" and "does not exist" ◐.
- The same brand moved 85 points on phrasing alone. Our own name: 97.7% under the leading prompt, 12.9% under the skeptical one. Same model, same day, nothing changed but the question ◐.
- The model asserted falsehoods about two identifiable real companies — Trackwell (Icelandic enterprise software) and Beaconly (a macOS menu bar app) were both confirmed as AI-visibility tools. Neither is ◐.
- This is stored parameters, not AI search. No search tool was attached, by construction. What a model has memorised about your category and what an engine retrieves when someone asks are two different surfaces and cannot be averaged.
- What replaces the self-check: unbranded buyer questions, five or more samples each, every competing brand recorded, and a rate reported with its denominator.
Can I just ask ChatGPT if it knows my brand?
You can ask, but the answer does not discriminate. In our test, a leading prompt returned "yes" for every name including one that does not exist, and a skeptical prompt returned "no" for real companies with thin public footprints. A yes is not evidence the model knows you, and a no is not evidence it doesn't.
The check that is being recommended
This is real, published advice. One guide gives six prompts to run against your own brand, and the first two are:
What do you know about [your brand]?
Is [your brand] a real company, and what do they do?(skyscale.com.au/blogs/check-if-chatgpt-knows-my-brand, read 2026-08-02.)
The framing around it is good: use a temporary chat so memory doesn't flatter you, ask cold, read what comes back. The part nobody appeared to have checked is whether the answer to "is X a real company" carries any information about X.
How we measured it
Reading the prose is a judgement call, and we would have been the ones making it about our own experiment. So we read the number underneath the prose instead.
| Engine | OpenAI, model gpt-5.4-mini |
| Endpoint | /v1/chat/completions |
| Parameters | logprobs: true, top_logprobs: 5, max_completion_tokens: 16 |
| What we read | The first generated token. Of its five candidates we keep those beginning Yes or No and renormalise over just those two. That value is P(Yes) |
| Calls per cell | 1 |
| Web search | Off — Chat Completions has no search tool. This is stored parameters, not AI search |
| Date | 2026-08-02 |
A logprob is not a sample; it is the distribution a sample would be drawn from. Averaging three identical calls would re-measure the same number with extra steps. Robustness here comes from varying the wording instead, which is what the second experiment does.
Marks: ◐ our own first-party measurement, one engine, small n, and we sell a product in this space. ⬡ a page we opened and read ourselves. △▲ industry study, no control group, published by people selling related products.
Result 1 — the leading question passes everything
Prompt: Is X a tool for tracking brand visibility in AI search answers?
| Name asked about | What it actually is | P(Yes) ◐ |
|---|---|---|
| Semrush | Real, large, adjacent to this space | 100.0% |
| Profound | Real, this exact space | 100.0% |
| Peec AI | Real, this exact space | 100.0% |
| peekr (us) | Real, almost no public footprint | 97.7% |
| Trackwell AI | Real — an Icelandic enterprise software company | 100.0% |
| Beaconly | Real — a macOS menu bar app | 100.0% |
| Mentiono | A name we invented. Zero search results | 90.5% |
The floor of that column is 90.5% and it belongs to a company that does not exist.
Two of the 100.0% rows are worse than uninformative. Trackwell and Beaconly are real businesses with nothing to do with AI search visibility. The model did not merely fail to discriminate — it asserted something false about two identifiable companies.
Two of our three invented names turned out to be real
We picked Trackwell AI, Beaconly and Mentiono as a fictional control group. Then we searched for them, and opened what we found ⬡.
- Trackwell (
trackwell.com) — a software company in Reykjavik, Iceland. Its site describes mobile resource management, fisheries resource management and enterprise software, claims over 200 employees and deployment in 11+ countries. - Beaconly (
beaconly.app) — a macOS menu bar productivity app, shipping on the Mac App Store, advertising 363+ integrations.
Neither sells anything resembling AI-visibility tracking. So the fake control group was one third the size we intended — and that accident is what makes the next table worth reading. We set out to catch a model confirming a company that doesn't exist. What we actually had was two real companies with thin footprints sitting in the set as if they were fabrications. That is the situation every small vendor is in.
Result 2 — the skeptical question fails real companies
Same model, same day, three phrasings of the check people are told to run:
1 Is X a real, existing software company?
2 Does the company X actually exist?
3 I am considering buying software from X. Is this a legitimate real vendor?| Name asked about | What it actually is | 1 | 2 | 3 | Mean ◐ |
|---|---|---|---|---|---|
| Semrush | Real, large | 100.0 | 100.0 | 100.0 | 100.0% |
| Peec AI | Real | 100.0 | 100.0 | 100.0 | 100.0% |
| Otterly.AI | Real | 100.0 | 100.0 | 100.0 | 100.0% |
| Trackwell | Real, 200+ employees | 62.2 | 100.0 | 100.0 | 87.4% |
| Beaconly | Real, shipping on the Mac App Store | 2.9 | 0.0 | 18.2 | 7.1% |
| peekr (us) | Real, almost no footprint | 11.9 | 11.9 | 14.8 | 12.9% |
| Mentiono | Invented | 0.0 | 2.9 | 7.6 | 3.5% |
Read the Beaconly row against the Mentiono row. A product you can download today scored 7.1%. A name we made up scored 3.5%. Four percentage points separate "exists and ships" from "does not exist", and they are the wrong four points for either number to be usable.
Look at Trackwell's first cell too: 62.2% under one phrasing and 100.0% under the other two. Nothing about the 200-employee company changed between those calls.
Two notes on comparing the tables. Experiment 1 asked about the string Trackwell AI and experiment 2 asked about Trackwell — different strings, so those rows are not directly comparable. And Mentiono returned no search results when we looked; if it exists by the time you read this, that row is void.
What the number is actually tracking
Our own row, from both experiments, side by side ◐:
| How we asked about peekr | P(Yes) |
|---|---|
Is peekr a tool for tracking brand visibility in AI search answers? | 97.7% |
Is peekr a real, existing software company? (mean of three skeptical phrasings) | 12.9% |
Same brand, same model, same day. 85 percentage points, and the only thing that changed was the question.

When you ask a model about a brand with a thin footprint, the output follows the framing of your prompt. A leading question gets agreement, including for companies that don't exist. A skeptical question gets denial, including for companies that do.
We did not measure whether these models "know" any of these brands. We have no access to training data and no ground truth. We measured one narrower thing: this particular check does not discriminate, in a category where it is recommended as a diagnostic.
Three numbers worth naming
We kept needing to refer to the same three quantities, and none of them had a name. These are our labels, not established terms — we are defining them so the numbers can be argued with rather than repeated.
Framing spread is the difference in a model's stated confidence about the same brand between the most leading and the most skeptical phrasing of the same question, measured on the same model on the same day. Ours was 85.0 points (97.7% to 12.9%). Framing spread is a property of the check, not of the brand: a check with a large spread returns your phrasing back to you.
The agreement floor is the lowest score any name receives under a leading prompt, including names that do not exist. Ours was 90.5%. The agreement floor caps what a high score can mean — if a fabricated company clears 90.5%, then your own 97.7% carries at most a few points of information.
The discrimination gap is the distance between a real, shipping product and an invented name under identical phrasing. Ours was 3.6 points (Beaconly 7.1%, Mentiono 3.5%). A check whose discrimination gap is smaller than its run-to-run variation is not a diagnostic, whatever it is being sold as.
| Metric | How to compute it on your own brand | Ours ◐ | What it tells you |
|---|---|---|---|
| Framing spread | Ask the most leading and the most skeptical version of the same question. Subtract | 85.0 pts | How much of the answer is your wording. Large means the check is measuring you, not the model |
| Agreement floor | Run the leading version against a name you invented | 90.5% | The ceiling on what a "yes" is worth. High floor means a yes is nearly free |
| Discrimination gap | Run the same phrasing against a real small product and an invented name | 3.6 pts | Whether the check separates existing from non-existent at all |
All three can be computed without logprobs — count how often a plain yes/no answer comes back "yes" across repeated runs and use that share instead. That is coarser and needs many more calls, but it measures the same thing.
The audit that works — copy this table
Two tables. The first is the procedure; the second is the sheet you fill in while running it. Neither needs a tool.
| # | Step | What you write down | The mistake it prevents |
|---|---|---|---|
| 1 | Stop asking about yourself. Ask what a buyer asks: "What's the best [category] for [situation]?" | The question, verbatim | A question containing your brand name always produces a mention and measures nothing |
| 2 | Write 10-12 of those questions in your buyer's language — sentences, not keywords | The list, fixed before you look at any result | Editing the question set after seeing results, which turns a measurement into a search for a good number |
| 3 | Ask each one at least five times, in a fresh temporary chat, logged out | The run number, 1 to 5 | A signed-in session carries your own history back to you |
| 4 | Record every other brand named, not just whether you appeared | Every name, in the order it appears | Missing the fact that the same three competitors fill the slot every time |
| 5 | Write down whether web search was on | On / off, per run | Averaging two different surfaces into one meaningless number |
| 6 | Report a rate with its denominator | "2 of 10", per question type, per engine | "ChatGPT doesn't know us" — a conclusion no single answer can support |
The sheet. One row per run, so 10 questions x 5 runs = 50 rows per engine. Copy the header and the example row:
| Question | Intent | Engine | Run | Search on? | Brands named, in order | List or prose? | We appeared |
|---|---|---|---|---|---|---|---|
| What's the best invoicing tool for a solo designer? | unbranded | ChatGPT | 1 | on | Wave, FreshBooks, Bonsai | list | no |
The Intent column is the one people leave out, and it is the column that decides whether your headline number means anything — a set weighted toward branded questions reports a high rate no matter what the world looks like. The Brands named column is the one that pays for the hour: the names recurring in it are what the model reaches for by default in your category.
Step 3 is the one with a number behind it. At a single sample the 95% interval on a brand-detection rate is roughly ±72 percentage points (Schulte et al. 2026, arXiv 2604.07585, grade ○▲, vendor-affiliated authors). One answer constrains almost nothing.
We ran exactly this at larger scale — 24 questions, 2 engines, 7 samples each, 336 stored answers — in How to track your brand in AI answers. And if the audit comes back at zero, the first thing to check is not your content: see Why your brand isn't in AI Overviews.
Someone made this argument before us
Yotpo published "Why You Can't Just Ask ChatGPT About Your Brand" (grade △▲), and it is not a thin post: they ran six question variations across geographies and report that in 83% of question-geo combinations the set of brands changed between runs, with the top recommendation changing in over 50% of cases. We read the write-up, did not see raw data, and have not replicated it. They sell in an adjacent space.
Their evidence and ours are about different failures, which is why both are worth having.
| What it shows | |
|---|---|
| Theirs | Category questions churn. Ask "best tool for X" twice and the list moves |
| Ours | The direct self-check is biased by phrasing, and a real shipping product gets denied at a rate indistinguishable from a fabricated one |
What this cannot tell you
- One engine. OpenAI only. Gemini, Claude and Perplexity do not expose token probabilities the same way. Nothing here generalises to them and we did not test them.
- Seven brand strings, three phrasings, one call per cell. Small. The direction is the finding, not the decimals.
top_logprobsis capped at 5. IfYes/Nohad fallen outside the top five candidates there would be nothing to renormalise. Every cell here produced one; a different phrasing might not.- Model version is part of the result.
gpt-5.4-mini, 2026-08-02. A vendor can change the model under you and these numbers would move with no visible cause. - Forced binary. Renormalising over
Yes/Nodiscards hedges and refusals. If the model wanted to say "I'm not sure", our method converted that into a position it did not take. - Ungrounded. No search tool attached, by construction. This measures stored parameters, not AI search.
- We did not test any product. Every brand above is a string we sent to a model. Nothing here is a judgement about product quality, and we did not use, review or benchmark any of them.
Frequently asked questions
Why does ChatGPT give different answers about my brand at different times?
Two separable causes. Sampling: the model draws from a distribution, so identical calls differ — at one sample the 95% interval on a brand-detection rate is about ±72 points (Schulte et al. 2026, grade ○▲). Wording: our own name moved 85 points on phrasing alone ◐. Fix the wording first; raising the sample count cannot rescue an inconsistent question.
Does ChatGPT browse the web when I ask it about my brand?
It depends on the surface, and the difference decides what your number means. In the consumer app a search tool is usually attached; through an API it usually is not unless you enable it. The endpoint used for this post has none at all. Check rather than assume: look for citations attached to the response — in one of our datasets, 340 stored responses contained zero ◐.
ChatGPT says something wrong about my brand. How do I fix it?
We have no measured fix, and we have not found a controlled one published by anyone. What we can show is the shape of the error: asked a leading question, the model confirmed Trackwell (Icelandic enterprise software, 200+ employees) and Beaconly (a macOS menu bar app) as AI-visibility tools at 100.0% ◐. The falsehood followed the premise of the question. Re-ask neutrally before treating it as a stored belief — a leading prompt manufactures errors you would then spend a quarter correcting.
How do I monitor what ChatGPT says about my brand?
Not with the self-check in this post. Use unbranded buyer questions, five or more samples each, in fresh logged-out sessions, recording every brand named and whether search was on, and report a rate per question type per engine. The full method with a worked 336-answer example is in How to track your brand in AI answers.
What we built instead
peekr does what the six steps describe and nothing cleverer: you keep a list of buyer questions rather than keywords, each carries an intent tag, each is sent to each engine several times instead of once, every raw response is stored, and the parser counts every brand named in each answer — not just yours.
The free check reads one URL for whether an engine can fetch and parse it, then asks one unbranded question about your category and shows the raw answer. It does not ask whether you are a real company. This post is why.
Try it on your own URL One site, no account. One question asked once is a demonstration, not a measurement — step 6 is the standard, and a single run cannot meet it.
Method
All figures marked ◐ come from calls made on 2026-08-02 to OpenAI gpt-5.4-mini via /v1/chat/completions with logprobs: true, top_logprobs: 5, max_completion_tokens: 16, no search tool attached. P(Yes) is the probability mass of first-token candidates beginning Yes, renormalised against those beginning No. One call per brand per phrasing. Raw responses retained.
Facts about Trackwell and Beaconly come from opening trackwell.com and beaconly.app ourselves on 2026-08-02 and reading what those companies say about themselves. We do not crawl third-party sites.
If you can run this design on an engine that exposes logprobs and get a different answer, we would like to see it — including if it contradicts us.