Should I target long-tail prompts instead of head terms for AI visibility? Only 31% of prompts trigger a search, and when one does the engine writes its own 5.48-word query — that is the one your page is matched against
The advice is everywhere and stated with unusual confidence: head terms are gone, AI Overviews eat them, and the way back in is long, specific, conversational queries. Search Engine Land ran it under the headline "Why AI optimization is just long-tail SEO done right." Several vendors sell exactly this analysis as their differentiator.
We measured one brand across both kinds of phrasing and got the opposite result, hard enough that we spent a while working out whether we had measured the wrong thing.
We had — and the reason turns out to be published, checkable and much more interesting than either position. There are two different query populations in this system, and the advice is about one while your retrieval is about the other.
Grading
| Mark | Meaning |
|---|---|
| ◐ | Our own first-party measurement. One brand, one category, one week, no control group, and we sell a product in this space. Maximum conflict of interest. |
| △▲ | Observational study, no control group, published by a company that sells a related product |
| ○▲ | Preprint with a statistical result, author affiliated with a vendor in this space |
| ✗ | Quoted widely, no method we could locate |
1. What our own data said, which looked like the opposite of the advice
One brand, one category (AI video clipping tools), July 2026. Same underlying question asked two ways, both engines, retrieval on, raw responses retained ◐:
| Natural-language question | Head-term phrasing | |
|---|---|---|
| e.g. | what's a good tool for pulling highlights out of long videos? | best AI clipping tools |
| Brand appeared in the answer | 0 / 120 | 3 / 12 |
| Brand's pages cited as a source | 0 / 242 | 7 / 154 |
| English + Gemini cell, page retrieved into context | 0 / 12 | 6 / 6, every time |
The head-term phrasing did not win narrowly. The natural-language phrasing produced nothing at all across 242 citations — on a brand whose page was pulled into context every single time when the same question was phrased as a head term.
So either the long-tail advice is wrong, or it is about something we did not test. It turns out to be the second, and the reason is a step in the middle that most of this discussion skips.
2. The step in the middle: the engine writes its own query
Nectiv, a marketing agency, analysed 8,500+ ChatGPT prompts across nine verticals (beauty, commerce, credit cards, fashion, jobs, local, software, real estate, travel) using their AI Tracker, extracting the 2,600+ search queries ChatGPT generated in response. Published 15 October 2025; written up in Search Engine Land. Grade △▲.
| Nectiv △▲ | |
|---|---|
| Prompts that triggered at least one search | 31% |
| Searches per prompt, when it searched | 2.17 |
| Average length of the query ChatGPT wrote | 5.48 words |
| Queries of five words or longer | 77% |
| Google average, for comparison (Semrush estimate) | ~3.4 words → ChatGPT's are 61% longer |
| Range by vertical | local 59% searched, credit cards 18%, fashion 19% |
Their stated limitations, in their words: "prompt tracking is already fuzzy," the study targeted commercial and buying intent specifically, and their n-gram analysis "skewed heavily towards the language I used in the prompts as well as the industries I decided to focus on." No data collection dates are given. We are quoting a study that tells you where it is soft, which is more than most in this area do.
Two things fall out of this that change the question.
(a) The prompt is not the query. People type long, contextual things into chatbots — that part of the long-tail story is real and we are not disputing it. But when a search happens, the model discards the phrasing and writes its own query, and that query averages under six words. Your page is retrieved against that, not against what the user typed.
(b) Most prompts in their sample did not search at all. 31% is a minority. The other ~69% were answered out of the model's parameters. ⚠ This is not a general statistic about all ChatGPT usage — their goal was commercial intent in nine verticals — but within that context, the majority of prompts were answered without anyone retrieving anything.
3. Our capture of the rewrite, which shows its shape
We recorded what the engines actually typed into a search box ◐:
Human asked: "AI 숏폼 클립 편집 툴 추천" (Korean, natural language)
ChatGPT searched: best AI short-form clip editing tools 2026 officialThat is a head term. The user's input was conversational; the query the engine generated is a keyword string with a year and the word official appended — the shape of a phrase somebody would type into Google, not the shape of the sentence the user wrote.
Two more things from the same capture, small samples both: ChatGPT rewrote 6 out of 6 Korean questions into English queries. Gemini kept 9 of 11 in Korean. ⚠ n = 6 and 11. Direction only.
This is what resolves section 1. Our natural-language questions produced zero citations not because long-tail is bad, but because a page optimised for a conversational sentence is being matched against a five-word rewrite. Nectiv measured the rewrite at scale; we caught one on camera. Independent evidence, same mechanism, and neither of us went looking for it.
4. "Long-tail" is three axes wearing one coat
With the rewrite in view, the popular advice can be pulled apart. It is claiming three things that vary independently:
| Axis | Two ends | Really about |
|---|---|---|
| A. Phrasing | keyword-shaped vs sentence-shaped | how the question is written — and who writes it |
| B. Specificity | best CRM vs best CRM for a two-person law firm | how many products are genuinely candidates |
| C. Brand presence | question names your brand vs describes a need | whether the model has to discover you |
"best invoicing software for freelance illustrators" is highly specific and still keyword-shaped. "what should I use for this?" is sentence-shaped and unspecific. The persona-page idea most people mean — accounting software for graphic designers — is a move along axis B, not axis A.
Sections 1 to 3 are an axis-A result and nothing more. And on axis A the answer is now clear enough to act on: you are not writing for the user's phrasing, you are writing for a five-word machine-generated query. Optimising a page to read like a conversation is optimising for a string that never reaches your page.
5. Axis B — untested by us, and the mechanism is what makes it plausible
This is where the persona-page idea lives, and we have no direct measurement. What we have is two observations that make the mechanism credible ◐:
There is room in the answer. Across our set, answers named 3–5 tools each, with 16–20 distinct tool names appearing across the whole set. The slots were not full. The brand was not crowded out; it was not a candidate.
Obscure names do get in. Among the domains cited were pexo.ai, socialync.io, quso.ai and reap.video — services with no meaningful brand recognition. There is published work showing visibility correlates strongly with brand prominence (Jack et al. 2026, ~37,000 runs, 533 brands, arXiv 2605.27439, grade △▲ — the authors sell AI-visibility consulting; even their lowest tier shows 3% exposure). Our observation does not contradict their numbers. It contradicts the story that unknown brands never reach the candidate set.
The mechanism, stated as a mechanism: if an answer has 3–5 slots and the candidate pool for best video editor is forty well-documented products, your odds are structural. If the pool for best video editor for a solo podcaster publishing weekly is six, the same slots are contested by far fewer names, and the prominence advantage that fills them at the head has less to work with.
🔴 We did not test this, and we have not found anyone who tested it with a control. Two figures circulate in support — long-tail earns 2x more AI Overview citations, 8+ word queries trigger overviews 7x more often. We could not trace either to a stated method, and we found the second stated two incompatible ways by different republishers, once as growth over time and once as a likelihood ratio. Grade ✗ for both.
⚠ Note also that Nectiv's rewrite finding cuts for axis B while cutting against axis A. A five-word machine query can still be a narrow one — best invoicing software freelance illustrators is five words. Specificity survives the rewrite. Conversational phrasing does not.
6. Axis C — and why the 31% makes our awkward dataset relevant
A separate dataset from the same brand: 336 stored answers, 24 questions, 2 engines, 7 samples each, split by whether the question contained the brand's own name ◐:
| Question type | Brand appeared |
|---|---|
| branded — the question names the brand | 100% |
| commercial — buying-intent questions | 57.1% |
| unbranded — the question describes a need | 0.0% |
| weighted total | 50.0% |
🔴 Retrieval was off for those 336 answers. They measure what the models had stored, not what AI search retrieves. We have always labelled this as a caveat.
★ Nectiv's 31% changes how much of a caveat it is. If, in commercial-intent contexts, roughly two-thirds of prompts are answered without a search, then the no-retrieval surface is not an artefact of our test rig — it is a large share of how the thing is actually used. What the model already believes about your category is doing most of the work, most of the time.
⚠ We are being careful with that inference: Nectiv measured nine verticals with a commercial-intent design, not general usage, and their tracking is self-described as fuzzy. But it moves this dataset from "a measurement we couldn't help taking" to "a measurement of a surface that matters," and we would rather say so than keep filing it under caveats.
And the practical point for persona pages: a persona question is unbranded by construction. It sits squarely in the cell where this brand scored zero out of every answer. Moving along axis B does not move you off axis C — it moves you further onto it.
So the persona-page proposal, honestly stated: a bet that specificity (axis B) buys enough advantage to overcome the cell (axis C) where you start at zero. That is a coherent bet. It is not a measured one.
7. What a persona comparison page has to survive
Say you write "The 3 best invoicing tools for freelance illustrators." Two of our measurements bear on what happens next, and both are cautions ◐.
When a vendor's own listicle gets cited, the vendor that published it appears in the answer 22% of the time. When a vendor's own product page is the cited source, it is 89%. In our set, 22 companies had their own article cited and their own name absent from the resulting answer. The mechanism is not mysterious: a "best tools" page is, to an answer engine, a list of candidate names with reasons attached, and eight or nine of those names belong to somebody else.
★ A persona top-3 page emits two competitor names instead of nine. We have not measured the top-3 or one-to-one genre and cannot tell you where it lands between 22% and 89%. The direction is plausible and untested, and publishing a number for it would be the exact thing this post complains about.
The ranking you assign yourself is not adopted. The brand ranked itself first in its own article. Reading that article, the engine placed it second once and fifth twice, called it good value every time, and gave "best overall" to a competitor in 19 of 21 answers.
⚠ A correction we owe on our own earlier reasoning: we first suspected engines were deliberately excluding a publisher from its own article. That was wrong — this brand survived at 50%, better than the 22% average. The variable is the genre of the page, not the identity of the publisher.
8. The precondition that makes all of this moot if you skip it
None of the above matters if the page does not rank.
Across 12 queries where Google showed an AI Overview, the brand's page was cited 6 out of 6 times when it held an organic position, and 0 out of 6 when it did not — Fisher's exact test, one-sided, p = 0.00108, zero exceptions in either direction ◐. Underneath it, a weaker pattern: the brand appeared as a recommendation in the overview's prose only on the two queries where it held first position. Positions 3, 8 and 10 produced citations but 0 of 4 recommendations (p = 0.067 — direction only, not significant).
This is the strongest single result we hold, and it points at organic rank rather than at prompt strategy. The version of the long-tail argument we find most defensible is not "AI prefers long-tail" — it is the much older "a narrow query is easier to rank for, and ranking is what feeds the answer." That requires no new mechanism, which is a point in its favour. And it now has a second leg under it: the query being ranked for is a five-word machine rewrite, which is a search-engine-shaped object.
⚠ n = 12, one brand, one category, one week. Direction worth believing; values not. Full write-up: Why isn't my brand showing up in AI search results?
9. How to test the persona bet before spending a quarter on it
- Write down which axis you are moving. Specificity, phrasing, or brand presence. A test that changes two at once cannot tell you which one did anything.
- Capture the rewrite, not just the prompt. Where the engine exposes the queries it generated, log them. That five-word string is your actual target keyword and you can read it directly rather than guessing — section 3.
- Count the candidate pool before you write. Ask the narrow question a few times and list every product named. If ten established names come back, the pool is not small and the premise is already false for that segment. One hour, and it can save the quarter.
- Use a control set of segments. Pick six persona segments; publish for three. The untouched three absorb background churn — roughly 65% of cited sources change between consecutive days (Schulte et al. 2026, arXiv 2604.07585, grade ○▲, vendor-affiliated authors).
- Measure rates, not sightings. At a single sample the 95% interval on a brand-detection rate is about ±72 percentage points (same source).
- Track whether the page ranks, as its own line. Given section 8, a persona page that never reaches the organic results for its query has not tested the persona hypothesis; it has tested nothing.
- Count who else appears in the answers your page gets cited in. That is the 22%-versus-89% question applied to your own page, and it is the number this genre most needs.
10. What this cannot tell you
- Axis B is untested by us. Sections 5 and 7 are mechanism and caution, not findings.
- One brand, one category, one week, across datasets with different sample sizes, one measured with retrieval off. We have labelled which is which and they should not be merged.
- We did not receive Nectiv's dataset, only their write-up and Search Engine Land's. Their sample is commercial-intent prompts in nine verticals with no stated collection dates, and does not generalise to all usage.
- The ~23-word figure for typical chatbot prompt length, which we have seen attributed to Semrush, we did not trace to a primary methodology page and have not relied on. The argument in section 2 rests on Nectiv's 5.48, which we did check.
- 🔴 No before-and-after with a control, on this or anything else — not from us, and not in anything we have read.
- The 89% figure may be reverse-caused — an engine could settle on a recommendation first and then visit that vendor's page to confirm it. Our data cannot separate those.
- We could not find published research on the persona-page genre specifically. If it exists, we would like the citation.
Method
Figures marked ◐ come from our own stored measurements on one brand in the AI video-editing category, July 2026, across three datasets reported separately and not merged: (a) engine responses via the OpenAI and Google APIs with retrieval enabled, across a natural-language query set (120 answers, 242 citations) and a head-term set (12 answers, 154 citations), including capture of the search queries the engines generated; (b) Google SERPs and AI Overviews via a SERP API across 13 queries, an overview appearing on 12; (c) 336 stored answers, 24 questions × 2 engines × 7 samples, with retrieval off. Raw responses retained. We do not crawl third-party sites. We did not test, review or benchmark any product named here, and the domains in section 5 are records of what was cited, not assessments.
External figures are attributed inline with publisher, date, method as stated, and a conflict-of-interest mark. Figures marked ✗ are ones we searched for and could not trace to a stated method; we listed rather than used them.
I build peekr, which is why these datasets existed to be pulled apart. Section 3 is the part I would keep if I could keep only one — we had that query capture sitting in our own database for weeks and read it as a curiosity about translation, until somebody else's study told us what it was.
If you have run section 9 on real persona segments, in either direction, I would like to read it.