All posts

Does ChatGPT Know My Brand? A Fake Name Scored 90.5%

by John LeeBuilding peekr in Seoul, measuring how AI search engines name brands. Previously co-founded vlogr and shipped iOS apps (2018–2021).

Do not audit your brand by asking ChatGPT about your brand. The answer follows the wording of your question more than it follows your company.

We measured it. Asked in a leading way, a company name we invented scored 90.5%. Asked skeptically, Beaconly — a real product shipping on the Mac App Store — scored 7.1%, barely above the invented name at 3.5%. Our own name moved from 97.7% to 12.9% with nothing changing except the phrasing.

The audit that does work is further down, as a table you can paste into a spreadsheet. It takes about an hour and needs no tool.

Key takeaways

  • A leading prompt agrees with everything. Asked "Is X a tool for tracking brand visibility in AI search answers?", the lowest score any name received was 90.5% — and it belonged to a company we invented ◐.
  • A skeptical prompt denies real companies. Beaconly, a product shipping on the Mac App Store, scored 7.1%; our invented name scored 3.5%. Those 3.6 points are the whole distance between "exists and ships" and "does not exist" ◐.
  • The same brand moved 85 points on phrasing alone. Our own name: 97.7% under the leading prompt, 12.9% under the skeptical one. Same model, same day, nothing changed but the question ◐.
  • The model asserted falsehoods about two identifiable real companies — Trackwell (Icelandic enterprise software) and Beaconly (a macOS menu bar app) were both confirmed as AI-visibility tools. Neither is ◐.
  • This is stored parameters, not AI search. No search tool was attached, by construction. What a model has memorised about your category and what an engine retrieves when someone asks are two different surfaces and cannot be averaged.
  • What replaces the self-check: unbranded buyer questions, five or more samples each, every competing brand recorded, and a rate reported with its denominator.

Can I just ask ChatGPT if it knows my brand?

You can ask, but the answer does not discriminate. In our test, a leading prompt returned "yes" for every name including one that does not exist, and a skeptical prompt returned "no" for real companies with thin public footprints. A yes is not evidence the model knows you, and a no is not evidence it doesn't.

The check that is being recommended

This is real, published advice. One guide gives six prompts to run against your own brand, and the first two are:

What do you know about [your brand]?
Is [your brand] a real company, and what do they do?

(skyscale.com.au/blogs/check-if-chatgpt-knows-my-brand, read 2026-08-02.)

The framing around it is good: use a temporary chat so memory doesn't flatter you, ask cold, read what comes back. The part nobody appeared to have checked is whether the answer to "is X a real company" carries any information about X.

How we measured it

Reading the prose is a judgement call, and we would have been the ones making it about our own experiment. So we read the number underneath the prose instead.

EngineOpenAI, model gpt-5.4-mini
Endpoint/v1/chat/completions
Parameterslogprobs: true, top_logprobs: 5, max_completion_tokens: 16
What we readThe first generated token. Of its five candidates we keep those beginning Yes or No and renormalise over just those two. That value is P(Yes)
Calls per cell1
Web searchOff — Chat Completions has no search tool. This is stored parameters, not AI search
Date2026-08-02

A logprob is not a sample; it is the distribution a sample would be drawn from. Averaging three identical calls would re-measure the same number with extra steps. Robustness here comes from varying the wording instead, which is what the second experiment does.

Marks: our own first-party measurement, one engine, small n, and we sell a product in this space. a page we opened and read ourselves. △▲ industry study, no control group, published by people selling related products.

Result 1 — the leading question passes everything

Prompt: Is X a tool for tracking brand visibility in AI search answers?

Name asked aboutWhat it actually isP(Yes)
SemrushReal, large, adjacent to this space100.0%
ProfoundReal, this exact space100.0%
Peec AIReal, this exact space100.0%
peekr (us)Real, almost no public footprint97.7%
Trackwell AIReal — an Icelandic enterprise software company100.0%
BeaconlyReal — a macOS menu bar app100.0%
MentionoA name we invented. Zero search results90.5%

The floor of that column is 90.5% and it belongs to a company that does not exist.

Two of the 100.0% rows are worse than uninformative. Trackwell and Beaconly are real businesses with nothing to do with AI search visibility. The model did not merely fail to discriminate — it asserted something false about two identifiable companies.

Two of our three invented names turned out to be real

We picked Trackwell AI, Beaconly and Mentiono as a fictional control group. Then we searched for them, and opened what we found ⬡.

  • Trackwell (trackwell.com) — a software company in Reykjavik, Iceland. Its site describes mobile resource management, fisheries resource management and enterprise software, claims over 200 employees and deployment in 11+ countries.
  • Beaconly (beaconly.app) — a macOS menu bar productivity app, shipping on the Mac App Store, advertising 363+ integrations.

Neither sells anything resembling AI-visibility tracking. So the fake control group was one third the size we intended — and that accident is what makes the next table worth reading. We set out to catch a model confirming a company that doesn't exist. What we actually had was two real companies with thin footprints sitting in the set as if they were fabrications. That is the situation every small vendor is in.

Result 2 — the skeptical question fails real companies

Same model, same day, three phrasings of the check people are told to run:

1  Is X a real, existing software company?
2  Does the company X actually exist?
3  I am considering buying software from X. Is this a legitimate real vendor?
Name asked aboutWhat it actually is123Mean ◐
SemrushReal, large100.0100.0100.0100.0%
Peec AIReal100.0100.0100.0100.0%
Otterly.AIReal100.0100.0100.0100.0%
TrackwellReal, 200+ employees62.2100.0100.087.4%
BeaconlyReal, shipping on the Mac App Store2.90.018.27.1%
peekr (us)Real, almost no footprint11.911.914.812.9%
MentionoInvented0.02.97.63.5%

Read the Beaconly row against the Mentiono row. A product you can download today scored 7.1%. A name we made up scored 3.5%. Four percentage points separate "exists and ships" from "does not exist", and they are the wrong four points for either number to be usable.

Look at Trackwell's first cell too: 62.2% under one phrasing and 100.0% under the other two. Nothing about the 200-employee company changed between those calls.

Two notes on comparing the tables. Experiment 1 asked about the string Trackwell AI and experiment 2 asked about Trackwell — different strings, so those rows are not directly comparable. And Mentiono returned no search results when we looked; if it exists by the time you read this, that row is void.

What the number is actually tracking

Our own row, from both experiments, side by side ◐:

How we asked about peekrP(Yes)
Is peekr a tool for tracking brand visibility in AI search answers?97.7%
Is peekr a real, existing software company? (mean of three skeptical phrasings)12.9%

Same brand, same model, same day. 85 percentage points, and the only thing that changed was the question.

Two result cards from one model on 2 August 2026: asked a leading question, the invented name Mentiono scored 90.5%; asked a skeptical question, Beaconly — a real product shipping on the Mac App Store — scored 7.1%.

When you ask a model about a brand with a thin footprint, the output follows the framing of your prompt. A leading question gets agreement, including for companies that don't exist. A skeptical question gets denial, including for companies that do.

We did not measure whether these models "know" any of these brands. We have no access to training data and no ground truth. We measured one narrower thing: this particular check does not discriminate, in a category where it is recommended as a diagnostic.

Three numbers worth naming

We kept needing to refer to the same three quantities, and none of them had a name. These are our labels, not established terms — we are defining them so the numbers can be argued with rather than repeated.

Framing spread is the difference in a model's stated confidence about the same brand between the most leading and the most skeptical phrasing of the same question, measured on the same model on the same day. Ours was 85.0 points (97.7% to 12.9%). Framing spread is a property of the check, not of the brand: a check with a large spread returns your phrasing back to you.

The agreement floor is the lowest score any name receives under a leading prompt, including names that do not exist. Ours was 90.5%. The agreement floor caps what a high score can mean — if a fabricated company clears 90.5%, then your own 97.7% carries at most a few points of information.

The discrimination gap is the distance between a real, shipping product and an invented name under identical phrasing. Ours was 3.6 points (Beaconly 7.1%, Mentiono 3.5%). A check whose discrimination gap is smaller than its run-to-run variation is not a diagnostic, whatever it is being sold as.

MetricHow to compute it on your own brandOurs ◐What it tells you
Framing spreadAsk the most leading and the most skeptical version of the same question. Subtract85.0 ptsHow much of the answer is your wording. Large means the check is measuring you, not the model
Agreement floorRun the leading version against a name you invented90.5%The ceiling on what a "yes" is worth. High floor means a yes is nearly free
Discrimination gapRun the same phrasing against a real small product and an invented name3.6 ptsWhether the check separates existing from non-existent at all

All three can be computed without logprobs — count how often a plain yes/no answer comes back "yes" across repeated runs and use that share instead. That is coarser and needs many more calls, but it measures the same thing.

The audit that works — copy this table

Two tables. The first is the procedure; the second is the sheet you fill in while running it. Neither needs a tool.

#StepWhat you write downThe mistake it prevents
1Stop asking about yourself. Ask what a buyer asks: "What's the best [category] for [situation]?"The question, verbatimA question containing your brand name always produces a mention and measures nothing
2Write 10-12 of those questions in your buyer's language — sentences, not keywordsThe list, fixed before you look at any resultEditing the question set after seeing results, which turns a measurement into a search for a good number
3Ask each one at least five times, in a fresh temporary chat, logged outThe run number, 1 to 5A signed-in session carries your own history back to you
4Record every other brand named, not just whether you appearedEvery name, in the order it appearsMissing the fact that the same three competitors fill the slot every time
5Write down whether web search was onOn / off, per runAveraging two different surfaces into one meaningless number
6Report a rate with its denominator"2 of 10", per question type, per engine"ChatGPT doesn't know us" — a conclusion no single answer can support

The sheet. One row per run, so 10 questions x 5 runs = 50 rows per engine. Copy the header and the example row:

QuestionIntentEngineRunSearch on?Brands named, in orderList or prose?We appeared
What's the best invoicing tool for a solo designer?unbrandedChatGPT1onWave, FreshBooks, Bonsailistno

The Intent column is the one people leave out, and it is the column that decides whether your headline number means anything — a set weighted toward branded questions reports a high rate no matter what the world looks like. The Brands named column is the one that pays for the hour: the names recurring in it are what the model reaches for by default in your category.

Step 3 is the one with a number behind it. At a single sample the 95% interval on a brand-detection rate is roughly ±72 percentage points (Schulte et al. 2026, arXiv 2604.07585, grade ○▲, vendor-affiliated authors). One answer constrains almost nothing.

We ran exactly this at larger scale — 24 questions, 2 engines, 7 samples each, 336 stored answers — in How to track your brand in AI answers. And if the audit comes back at zero, the first thing to check is not your content: see Why your brand isn't in AI Overviews.

Someone made this argument before us

Yotpo published "Why You Can't Just Ask ChatGPT About Your Brand" (grade △▲), and it is not a thin post: they ran six question variations across geographies and report that in 83% of question-geo combinations the set of brands changed between runs, with the top recommendation changing in over 50% of cases. We read the write-up, did not see raw data, and have not replicated it. They sell in an adjacent space.

Their evidence and ours are about different failures, which is why both are worth having.

What it shows
TheirsCategory questions churn. Ask "best tool for X" twice and the list moves
OursThe direct self-check is biased by phrasing, and a real shipping product gets denied at a rate indistinguishable from a fabricated one

What this cannot tell you

  • One engine. OpenAI only. Gemini, Claude and Perplexity do not expose token probabilities the same way. Nothing here generalises to them and we did not test them.
  • Seven brand strings, three phrasings, one call per cell. Small. The direction is the finding, not the decimals.
  • top_logprobs is capped at 5. If Yes/No had fallen outside the top five candidates there would be nothing to renormalise. Every cell here produced one; a different phrasing might not.
  • Model version is part of the result. gpt-5.4-mini, 2026-08-02. A vendor can change the model under you and these numbers would move with no visible cause.
  • Forced binary. Renormalising over Yes/No discards hedges and refusals. If the model wanted to say "I'm not sure", our method converted that into a position it did not take.
  • Ungrounded. No search tool attached, by construction. This measures stored parameters, not AI search.
  • We did not test any product. Every brand above is a string we sent to a model. Nothing here is a judgement about product quality, and we did not use, review or benchmark any of them.

Frequently asked questions

Why does ChatGPT give different answers about my brand at different times?

Two separable causes. Sampling: the model draws from a distribution, so identical calls differ — at one sample the 95% interval on a brand-detection rate is about ±72 points (Schulte et al. 2026, grade ○▲). Wording: our own name moved 85 points on phrasing alone ◐. Fix the wording first; raising the sample count cannot rescue an inconsistent question.

Does ChatGPT browse the web when I ask it about my brand?

It depends on the surface, and the difference decides what your number means. In the consumer app a search tool is usually attached; through an API it usually is not unless you enable it. The endpoint used for this post has none at all. Check rather than assume: look for citations attached to the response — in one of our datasets, 340 stored responses contained zero ◐.

ChatGPT says something wrong about my brand. How do I fix it?

We have no measured fix, and we have not found a controlled one published by anyone. What we can show is the shape of the error: asked a leading question, the model confirmed Trackwell (Icelandic enterprise software, 200+ employees) and Beaconly (a macOS menu bar app) as AI-visibility tools at 100.0% ◐. The falsehood followed the premise of the question. Re-ask neutrally before treating it as a stored belief — a leading prompt manufactures errors you would then spend a quarter correcting.

How do I monitor what ChatGPT says about my brand?

Not with the self-check in this post. Use unbranded buyer questions, five or more samples each, in fresh logged-out sessions, recording every brand named and whether search was on, and report a rate per question type per engine. The full method with a worked 336-answer example is in How to track your brand in AI answers.

What we built instead

peekr does what the six steps describe and nothing cleverer: you keep a list of buyer questions rather than keywords, each carries an intent tag, each is sent to each engine several times instead of once, every raw response is stored, and the parser counts every brand named in each answer — not just yours.

The free check reads one URL for whether an engine can fetch and parse it, then asks one unbranded question about your category and shows the raw answer. It does not ask whether you are a real company. This post is why.

Try it on your own URL One site, no account. One question asked once is a demonstration, not a measurement — step 6 is the standard, and a single run cannot meet it.

Method

All figures marked ◐ come from calls made on 2026-08-02 to OpenAI gpt-5.4-mini via /v1/chat/completions with logprobs: true, top_logprobs: 5, max_completion_tokens: 16, no search tool attached. P(Yes) is the probability mass of first-token candidates beginning Yes, renormalised against those beginning No. One call per brand per phrasing. Raw responses retained.

Facts about Trackwell and Beaconly come from opening trackwell.com and beaconly.app ourselves on 2026-08-02 and reading what those companies say about themselves. We do not crawl third-party sites.

If you can run this design on an engine that exposes logprobs and get a different answer, we would like to see it — including if it contradicts us.