Why isn't my brand showing up in AI search results? In our test it came down to Google organic rank (6/6 vs 0/6)
Most writing about getting cited by AI is about the page. Add a schema block. Add statistics. Add a quote. Add an FAQ section. Make it "AI-readable."
We did all of that, measured it, and got zero.
Then we went looking for what actually moved. This post is the log of four experiments — including the two that failed and the hypothesis of ours that our own data killed. The sample is small and we say so in every section. The direction is what we'd ask you to take away, not the values.
Grading
We mark every number so you can tell what it's worth:
| Mark | Meaning |
|---|---|
| ◎ | Peer-reviewed, control group |
| ○ | Preprint with a control group or statistical test |
| △ | Preprint, observational, no control |
| ▲ | Author sells a related product (conflict of interest) |
| ◐ | Our own first-party measurement — no peer review, no control group, tiny n, and we sell a product in this space. Maximum conflict of interest. |
Everything marked ◐ below is ours. Read it as "we observed this, under these conditions", never as "AI search works like this."
The setup
One brand, one category (AI video clipping tools), one measurement window in July 2026. We queried Google SERPs and Google AI Overviews through a SERP API, and ChatGPT / Gemini through their APIs, with the raw responses retained.
The brand's site is healthy by every conventional measure: ~1,290 pages indexed, blog posts holding Google positions 1, 3, 8 and 10 for head terms in its category, tables, first-party numbers, author bylines, schema markup.
Its appearance rate in non-branded AI answers was 0.0%.
Experiment 1 — Is it the language, or the grounding? No.
Hypothesis: the brand is invisible in AI answers because we were asking in Korean, or because the model wasn't retrieving at all.
We ran a 2×2: {English, Korean} × {grounded, ungrounded}, natural-language questions, both engines, samples retained.
Result ◐: 0 appearances out of 36. All four cells zero. Language wasn't the discriminator, and neither was grounding — turning retrieval on didn't produce a single mention.
This matters because it cuts against a claim we had expected to lean on. The literature says querying in the local language substantially raises local champions' recommendation share (+0.80 vs +0.15 for global multinationals, t = −8.84, p < 0.001 — Żatuchin et al. 2026, Language Blind Spot, arXiv 2606.23165; grade ○▲, one author is affiliated with a vendor in this space).
That study's sample was 66 known brands across 11 markets. It doesn't contain unknown brands. Our reading now: the language effect applies to brands the model already has a representation of. If the model has no candidate to promote, switching languages promotes nothing. That's a scope condition, not a refutation.
Experiment 2 — Is it how the question is phrased? Partly, and it revealed something else.
Hypothesis: natural-language questions ("what's a good tool for pulling highlights out of long videos?") behave differently from search-style head terms ("best AI clipping tools").
Result ◐: they do, and much more on the citation axis than on the appearance axis.
| Natural-language question | Head-term phrasing | |
|---|---|---|
| Brand appeared in answer | 0 / 120 | 3 / 12 |
| Brand's pages cited as a source | 0 / 242 | 7 / 154 |
| English + Gemini only | 0 / 12 | 6 / 6, every time |
Two things fell out of this that we did not go looking for.
(a) Retrieval and selection are different failures. In the English/Gemini/head-term cell the brand's page was retrieved 6/6 — every single time — and the brand still appeared in only 3/12 answers. The document was in the context window and was not chosen. Anyone optimizing content when their actual failure is at selection is optimizing the wrong stage. Anyone optimizing content when their failure is at retrieval is doing something that can't work at all.
(b) URL-level targeting doesn't hold. We aimed a query at a specific article. The engine retrieved a different article from the same blog. You do not get to choose which of your pages gets pulled in.
We also captured what the engines actually typed into a search box, which is where the third experiment came from:
Human asked: "AI 숏폼 클립 편집 툴 추천" (Korean)
ChatGPT searched: best AI short-form clip editing tools 2026 official◐ ChatGPT rewrote 6 out of 6 Korean questions into English search queries. Gemini kept 9 of 11 in Korean. And ChatGPT appends the word official — it is deliberately hunting for vendor product pages, not reviews. (Sample: 6 and 11 queries. Tiny. Direction only.)
Experiment 3 — What actually predicts citation? Organic rank.
This is the one that changed what we do.
We took 13 queries, pulled the Google SERP and the AI Overview for each, and asked one question: when the AI Overview appeared, did it cite the brand — and was the brand ranking organically for that same query?
AI Overviews appeared on 12 of the 13 queries.
| Of the 12 queries with an AI Overview | Brand cited in the AI Overview |
|---|---|
| Brand held an organic position (6) | 6 / 6 |
| Brand held no organic position (6) | 0 / 6 |
Fisher's exact test, one-sided: p = 0.00108. Zero exceptions in either direction. ◐
And a second, weaker pattern underneath it: the brand appeared as a recommendation in the AI Overview's prose only on the two queries where it ranked #1. Positions 3, 8 and 10 produced citations but 0/4 recommendations (p = 0.067 — direction only, not significant).
The five hypotheses this killed
Before running this we had five competing explanations. We can now rule out four of them for this brand, in this category, in this window:
| Hypothesis | Verdict ◐ | Evidence |
|---|---|---|
| Indexing — Google doesn't know the pages | ❌ Not it | Exact-title query returns the target page at #1. ~1,290 pages indexed. |
| Country — US vs KR search results differ | ❌ Not it | The same English query returned identical positions from both locales (1/1, 3/3). |
| Language | ❌ Not it | Experiment 1: all four language×grounding cells were zero. |
| Fame — you have to be a known brand | ❌ Not it | See below. |
| Slot scarcity — the answer is full | ❌ Not it | See below. |
| Organic rank | ✅ This one | The table above. |
On fame ◐: among the domains actually cited in these answers were pexo.ai, socialync.io, quso.ai, reap.video — services with no meaningful brand recognition. Obscure tools get cited. There is published work showing visibility correlates strongly with brand prominence (Jack et al. 2026, ~37,000 runs, 533 brands, arXiv 2605.27439 — grade △▲, the authors sell AI-visibility consulting; their lowest tier still shows 3% exposure). Our observation doesn't contradict their numbers. It contradicts the narrative that low-prominence brands don't reach the candidate set at all. They do.
On slot scarcity ◐: answers named 3–5 tools each, and 16–20 distinct tool names appeared across the set. There was room. The brand simply wasn't a candidate.
Experiment 4 — Can we see Bing? No.
Hypothesis: ChatGPT's search is Bing-based, so Bing rank should predict ChatGPT citation the way Google rank predicts AI Overview citation.
Result: we failed. We could not obtain Bing SERP data we trust. Worse, we found that our SERP provider returns status_code: 20000, "Ok." while handing back a Bing result set that is silently wrong.
So the ChatGPT half of the mechanism is unconfirmed. We know ChatGPT hunts for official vendor product pages. We do not know what ranking surface decides which ones it finds. If anyone has measured Bing rank against ChatGPT citation with a control, we want to read it.
What we think this means
Two mechanisms, not one.
| Engine | What it appears to read ◐ | So the work is |
|---|---|---|
| Google AI Overview / Gemini | Rides Google organic. Vendor-owned pages 52%, YouTube 27%, community 8.5% of AI Overview citations (n = 154 citations) | Conventional organic rank. Not a separate discipline. |
| ChatGPT | Vendor product/about pages 66%, review sites like G2/Capterra 19%, tech media 13%, vendor listicles 3% | Getting the product page — not the blog — into whatever surface it searches |
If you run one playbook against both, you will fix one and leave the other flat, and your dashboard will average the two into something meaningless.
The literature and our data agree more than they look like they do. The most-cited GEO result — Aggarwal et al. 2024 (Princeton, arXiv 2311.09735, grade ○) — reports large gains from adding quotations (+42.6%) and statistics (+32.8%). That experiment measures a document already sitting in a five-document context. It is a conditional-on-retrieval effect. The peer-reviewed replication attempt (Puerto et al. 2025, C-SEO Bench, NeurIPS D&B 2025, arXiv 2506.11097, grade ◎) found 3 significant positive results out of 54 technique×domain combinations, and 0 on QA tasks — and found that the gains collapse toward zero as more competitors adopt the same technique (independently reproduced by Chu & Hou 2026: +0.802 for a lone adopter → +0.007 when all nine adopt).
The 45-study survey (arXiv 2607.14035) puts it flatly: "no technique with a stable, longitudinal, cross-platform causal effect."
Our ◐ finding is the same shape from the other end. We had every content attribute the Princeton paper recommends, and got 0/120 — because we were failing at a stage those experiments hold constant.
What we still don't know
This is the section we'd most like other people to attack.
- n = 12. The 6/6 vs 0/6 split is clean and the p-value is real, but it is twelve queries, one brand, one category, one week. The direction is worth believing. The values are not.
- We measured each AI Overview once. We enforce repeat sampling on the LLM side because single measurements of generative engines are close to uninformative (Schulte et al. 2026, arXiv 2604.07585, grade ○▲: at n=1 the 95% CI on brand detection is ±0.724; ~65% of cited sources change between consecutive days). We did not hold ourselves to that standard on the AI Overview side. That's a defect in our own method, not a footnote.
- Correlation, not proof. "Ranks organically" and "gets cited" could both be downstream of something we didn't measure. We have no intervention data — nobody has shown us a before/after where a page moved up and citations followed.
- Which ChatGPT surface. Unresolved, see Experiment 4.
- Why a sibling page outranks the targeted one. We can't inspect it. Opening the page ourselves would make us a crawler, and we've ruled that out.
- How often to re-measure. Nobody knows how long an intervention takes to show up. We don't either.
- Korean. We measured Korean nine times and never saw the brand. The honest statement is not "0% in Korean" — it's "nine measurements can't tell you." Separately: this brand had zero `/ko` pages indexed, while Czech, Greek, Hungarian and Romanian all existed. For Korean this isn't a ranking problem yet. There's nothing entered in the race.
Two things we'd do differently tomorrow
Find out which stage you're failing at before you optimize anything. Retrieved-but-not-selected and never-retrieved look identical on a dashboard and need opposite responses. Ours were both, on different engines, at the same time.
Check whether your own "best tools" post is working for you. We measured what happens when a vendor-published listicle gets cited: the publishing company itself appears in the answer 22% of the time, versus 89% when its product page is the cited source ◐. In our set, 22 companies had their own article cited and their own name absent from the answer.
One brand ranked itself #1 in its own article. The AI read that article and placed it 2nd once and 5th twice, described it as "good value" every time, and gave "best overall" to a competitor in 19 of 21 answers. Self-assigned rank is not adopted.
⚠ Two cautions on those numbers. The 89% may be reverse-caused — the engine could decide on a recommendation first and then visit that vendor's homepage to confirm it. Our data can't separate those. And we initially suspected the engine was deliberately excluding the brand from its own article. That was wrong; the brand survived at 50%, better than the 22% average. The variable is the genre of the page, not the identity of the publisher.
We looked for published research on this specific effect — vendor-published listicles and who benefits — and found none. If you know of any, we'd like the citation.
Method and reproduction
Queries, engines, dates, sample counts and raw response handling are described above per experiment. We keep raw responses. We do not crawl third-party sites. Total spend on the rank experiment was under $0.25.
We build peekr, which is how we had this data — it samples the same question multiple times per engine and keeps the raw responses, because single measurements of generative engines don't mean much. That's the whole product pitch and we'll stop there; the findings above stand or fall on their own method, and we'd rather you argue with the method.
If you can replicate this on a different brand or a different category, please publish it — including if you get the opposite result. There are, as far as we can find, zero published studies measuring AI visibility for Korean-language brands. We'd like that number to stop being zero.