Does YouTube help my brand show up in AI answers? Published figures for the same engine range from 0.2% to 26% — the denominators are different, and almost nobody says which one they are quoting
Somebody has almost certainly told you to start a YouTube channel because AI cites YouTube.
The advice is not baseless. There is more real data behind it than behind most GEO advice, including one study with a hundred million citation instances in it. The problem is that the numbers people quote from those studies are not measuring the same fraction, and when they get lined up next to each other they produce apparent contradictions of two orders of magnitude.
This post lines them up properly, adds our own much smaller measurement to the same table, and separates the part that is well established from the part that is being sold.
Grading
| Mark | Meaning |
|---|---|
| ◐ | Our own first-party measurement. No peer review, no control group, one brand, one category, one week, and we sell a product in this space. Maximum conflict of interest. |
| △▲ | Observational study, no control group, published by a company that sells a related product |
| ○▲ | Preprint with a statistical result, author affiliated with a vendor in this space |
Everything below is △▲ or ◐. There is no peer-reviewed work on this question that we have been able to find.
1. Three claims that keep getting compressed into one
- AI answers cite YouTube a lot. — Well supported. Multiple independent datasets, including ours.
- *AI answers cite YouTube a lot in general.* — False as stated, and every dataset that reports an engine breakdown shows why.
- Therefore publishing YouTube videos will get your brand into AI answers. — Not shown by any of these studies, and the largest of them says so explicitly.
Claim 1 is the finding. Claim 3 is the pitch. Claim 2 is the join, and it is where the arithmetic goes wrong.
2. What is actually published
BrightEdge, via its AI Catalyst platform, analysing YouTube citation patterns across Google AI Overviews, Google AI Mode, ChatGPT and Perplexity from May 2024 to September 2025, written up in Search Engine Land in January 2026. Grade △▲ — BrightEdge sells enterprise SEO software; no sample size, confidence interval or limitation is stated anywhere in the write-up.
| BrightEdge △▲ | YouTube share |
|---|---|
| Google AI Overviews | 29.5% |
| Google AI Mode | 16.6% |
| Perplexity | 9.7% |
| quoted average across platforms | 20% |
They also report YouTube cited 200 times more than any other video platform, and note ChatGPT's YouTube citations growing "off a small base."
Otterly.AI, YouTube AI Citation Study 2026: more than 100 million AI citation instances over a 30-day period, across six platforms — ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Copilot and Gemini. They identified YouTube videos already cited by these systems, pulled metadata (views, likes, duration, title and description length, subscriber count, publish date, format), and ran Pearson correlations. Grade △▲ — Otterly sells an AI visibility tool.
| Otterly △▲ | Figure |
|---|---|
| Perplexity | 38.7% |
| Google AI Overviews | 36.6% |
| Copilot | 0.5% |
| Gemini | 0.2% |
| Long-form video share of citations | 94% (Shorts 5.7%, playlists and livestreams 0.3%) |
| Views, likes, subscriber counts | r ≈ −0.03 to 0.02 |
| Description length | r ≈ 0.31 |
| Hashtag presence | r ≈ 0.20 |
★ Credit where it is due: this is a more careful study than the genre norm. The authors state that correlation does not imply causation, and they state the limitation that decides the whole question — their dataset contains only videos that were already cited, so the findings are "strongest for explaining repeated citation behavior, not initial eligibility." Almost nobody quoting them repeats that sentence. It is the most important sentence in the study.
3. 🔴 Why these two tables look like they contradict each other, and do not
Perplexity is 9.7% in one and 38.7% in the other. Gemini is 0.2% in one and, in our data, 26%.
Read Otterly's own framing again: Perplexity and AI Overviews "drove most YouTube citations." That is share of YouTube citations, split by which platform produced them. BrightEdge's is share of each platform's citations that happened to be YouTube.
| Denominator | Question it answers | |
|---|---|---|
| BrightEdge style | all citations made by platform X | When Perplexity cites something, how often is it YouTube? |
| Otterly style | all YouTube citations, all platforms | Of the YouTube citations out there, which platform made them? |
Both are legitimate. They are not comparable, and only one of them tells you where to spend a video budget. Otterly's Gemini figure of 0.2% is compatible with Gemini citing YouTube constantly inside its own answers, if Gemini contributed few citations overall to their 30-day collection. It is not evidence that Gemini ignores video.
⚠ We inferred the denominator rather than being told it. Their platform figures sum toward 100% and their phrasing is share-of-YouTube-citations language. That is an inference from two clues, not a statement we were given, and if Otterly tells us we have read it wrong we will correct this section and say so here.
★ This is the practical takeaway of the post. The next time a deck shows you a YouTube citation share, the only question worth asking is percent of what — and in our experience of reading this literature, roughly half the time nobody in the room knows.
4. Our own numbers, held to the same standard
One brand, one category (AI video clipping tools), July 2026, raw responses retained. We were not studying YouTube; we were counting what got cited, and YouTube kept showing up unevenly. All the figures below use the BrightEdge-style denominator — share of that surface's citations.
| Surface ◐ | YouTube share | Denominator |
|---|---|---|
| Google AI Overviews | 26.6% | citations attached to the overviews on 12 of 13 queries |
| LLM grounding, natural-language questions, both engines pooled | 13% | 242 citations across 72 distinct domains |
| LLM grounding, head terms — ChatGPT | 0% | 154 citations, split by engine |
| LLM grounding, head terms — Gemini | 26% | same 154 |
| Google AI Overviews, Korean-language queries only | 32% | 25 citations across 4 queries |
🔴 Those rows come from four different query sets and must not be averaged. Committing that error while describing it would be an unusually stupid way to write this post.
⚠ Our AI Overview figure and BrightEdge's are close — 26.6% against 29.5%, from sample sizes that differ by orders of magnitude. We are not calling that replication; two numbers agreeing once proves very little, and we would have reached for an explanation if they had disagreed. We mention it because we would have mentioned a disagreement.
On the Korean row. The brand we measured had zero Korean-language pages indexed while fourteen other locales — Czech, Greek, Hungarian, Romanian among them — were indexed. On the Korean Gemini surface the cited properties were concrete and countable: YouTube 12, Tistory 8, a Korean tech outlet 3. In that market the brand was not losing a ranking contest. It had not entered one. ⚠ Four queries. Direction only.
5. What all the data agrees on
Strip out the denominator confusion and three independent datasets support the same two statements:
- On Google's answer surfaces, YouTube is a large minority of citations. BrightEdge 29.5%, Otterly 36.6% of YouTube citations arriving from that surface, us 26.6%. Whatever else is true, this one is not a single vendor's artefact.
- The engine spread is enormous, and it is the thing that should drive your budget. In our head-term set, ChatGPT cited vendor product and about pages 66% of the time and review sites 19%, and video 0%. Gemini in the same set cited video 26% and review sites 0%. BrightEdge's own note that ChatGPT is growing "off a small base" points the same way.
One playbook run against both engines will move one and leave the other flat, and a dashboard that averages them will show a number describing neither.
6. The mechanism, and the part of it that is genuinely measured
There is a plausible story: a YouTube page is not video to a retrieval system. It is a title, a description, a chapter list and a full timed transcript, on a domain with extreme crawl priority, in a format where a specific timestamp can be pointed at as the answer to a specific question.
Otterly's correlations are the best available evidence on this and they are worth taking seriously, with their own caveat attached:
- Popularity does essentially nothing. Views, likes and subscriber counts land at r ≈ −0.03 to 0.02 — that is noise. If citation were a popularity contest, this is where it would show.
- Text does something. Description length reaches r ≈ 0.31, hashtags r ≈ 0.20. Weak, but on the text axis rather than the audience axis, which is what the mechanism predicts.
- Long-form takes 94% of citations against 5.7% for Shorts. A longer transcript is more surface to match against and more to time-stamp.
⚠ All three are correlations within already-cited videos, which is the limitation Otterly states. They describe what distinguishes a repeatedly-cited video from a rarely-cited one. They do not describe what gets a video cited for the first time.
How we treat this internally: we have no video, so we have run none of it. What we took from it is a change to what we would count — if audience signals are noise and text signals are not, a video asset should be measured as a text page with a transcript, not as a channel with subscribers. That is a cheap change to make before you have any data of your own.
7. 🔴 The step nobody measured
| Question | Who has measured it |
|---|---|
| Do AI answers cite YouTube? | BrightEdge △▲, Otterly △▲, us ◐ |
| Does the share differ by engine? | All three — and all three say yes, strongly |
| Among already-cited videos, what distinguishes the ones cited most? | Otterly △▲, with correlations and a stated caveat |
| If I publish videos, does my brand's appearance rate rise? | 🔴 Nobody. Otterly says this directly: their data speaks to repeated citation, "not initial eligibility." |
Everything written under headings like how to get your videos cited by AI is answering the fourth row using evidence from the first three. That may still turn out to be correct. It is not currently shown, and the largest study in the field is the one telling you so.
Our own position is worse, not better: the brand we measured owned no video assets at all, so we never observed a brand crossing that line either. We can tell you how big the slot is. We cannot tell you what it costs to get into it.
8. How to test it on your own brand
- Establish a before, as a rate rather than a sighting. Pick 10–15 questions your buyers would type; ask each repeatedly. At a single sample the 95% interval on a brand-detection rate is roughly ±72 percentage points (Schulte et al. 2026, arXiv 2604.07585, grade ○▲ — vendor-affiliated authors). Record per engine, never pooled — sections 3 and 5 are the reason.
- Split into treatment and control halves. Publish video targeting only the treatment half. The control half is what tells you whether movement was yours or the engine's.
- Publish so that a null result means something. Full transcript, chapters, long-form, a title that is the question a person would type. Section 6 says the transcript and description are the load-bearing parts. If you publish uncaptioned footage and nothing happens, you have tested a different hypothesis.
- Fix the observation window in advance. Nobody has published how long an intervention takes to reach these surfaces. We do not know either. Choosing the window afterwards is how a null becomes a positive.
- Re-measure identically, including the parts you would rather change.
- Report the control half, especially if it moved. Roughly 65% of cited sources change between consecutive days in published measurement (Schulte et al. 2026, ○▲). A surface churning that fast will produce movement in a group that received no treatment.
If step 6 shows your control moving as much as your treatment, you have learned the most valuable thing available here: that you cannot yet attribute anything.
9. What this post cannot tell you
- Our figures are one brand, one category, one week, across four query sets that are not directly comparable to each other.
- We did not measure Perplexity or Copilot at all. Those empty cells are facts about us.
- We measured each AI Overview once. We enforce repeat sampling on the LLM side and did not hold ourselves to it here. That is a defect in the method, not a footnote.
- We did not receive either external dataset. We read both write-ups; neither underlying dataset appears to be public.
- The denominator reading in section 3 is our inference, from phrasing and from figures that sum toward 100%. We would like to be corrected if it is wrong.
- 🔴 We cannot tell you that publishing video will change your numbers, and neither can the studies. No before-and-after with a control exists in anything we have read.
Method
External figures are attributed inline with publisher, date, method as stated by the publisher, and a conflict-of-interest mark. We read the Search Engine Land write-up of BrightEdge's analysis and Otterly.AI's study page directly. We did not inspect either underlying dataset.
Figures marked ◐ come from our own stored measurements on one brand in the AI video-editing category, July 2026: Google SERPs and AI Overviews retrieved through a SERP API across 13 queries, an overview appearing on 12; and engine responses via the OpenAI and Google APIs with retrieval enabled, across a natural-language query set producing 242 citations over 72 domains and a head-term set producing 154 citations. Raw responses retained. We do not crawl third-party sites, and we did not open, watch, test or review any video or product named or implied here.
I build peekr, which is why our citation data existed to be counted. Section 3 is the part I would keep if I could keep only one — the denominator question is not a technicality, it is the difference between a number you can act on and a number you can only repeat.
If you run the section 8 design and get a result in either direction, I would like to read it.