← All studies

Study

YouTube in AI answers: why Gemini and Perplexity quote videos, and what they actually read

By Simon Vasconcelos Lee

Between July 10 and 24, 2026, we archived 883 answers from ChatGPT, Claude, Gemini and Perplexity, asked verbatim, on a weekly cadence, across 69 questions from three markets that have nothing in common. On that corpus, youtube.com is the most cited source on Gemini and the most cited source on Perplexity, all domains included. On the other two engines, it does not appear once in 1,655 recorded citations. This study takes that asymmetry apart: which engines quote videos, what part of a video they actually read, and what that makes worth doing.

What was measured

The starting point was a map, not a hunch. On the Atlas, the registry of sources our surveys surface, youtube.com sits at an AI Authority of 100 on Gemini and 100 on Perplexity: on a 0 to 100 scale computed across every measurement Epovest runs over the last 30 days, it is the source those two engines cite the most. The two other columns are empty. A map that surprising deserves to be checked against the ground, so we opened the full archive behind it.

The corpus is the complete answer history of our own trackers: 16 weekly trackers, 69 distinct questions, asked word for word through the engines' official interfaces, on three unrelated fronts: travel destinations in English, AI visibility in English and French, web services in French. Every answer is archived with its cited sources, in order, dated. Two weeks of surveys give 883 answers and 6,754 recorded citations.

Engine Answers Answers citing sources Citations recorded YouTube citations Answers citing YouTube
ChatGPT 223 159 1,213 0 0
Claude 223 87 442 0 0
Gemini 223 211 2,546 35 35, or 16.6% of sourced answers
Perplexity 214 214 2,553 205 127, or 59.3% of sourced answers

The asymmetry is the finding. Whether a video can carry a name into an AI answer is not a property of YouTube: it is a property of each engine's pipeline.

One domain, four different treatments

Perplexity quotes it constantly. More than half of its sourced answers cite at least one video: 127 answers out of 214, with 205 video citations in total, 8% of everything the engine cited over the two weeks. Ten times, a video opened the source list. The average position, 8.3 on source lists of about twelve, says the rest: videos are rarely the headline source, they are a permanent presence in the supporting cast.

Gemini quotes it at the top of a very long tail. 35 citations put youtube.com first among the hundreds of domains Gemini cited, just ahead of reddit.com at 33. That is only 1.4% of its citations, spread across 16.6% of its sourced answers, but no other single domain did better, and twice a video was the first source of the answer. The presence covers all three fronts: 22 citations on travel questions, 11 on AI visibility, 2 on web services.

Claude reads it and writes without it. Zero citations in 442. But the archive shows the engine fetching a YouTube watch page five times across the fortnight, two distinct videos, each carrying the question nearly verbatim in its title. The pages were retrieved; the answers were written without them.

ChatGPT does not touch it. Zero citations in 1,213, and not one YouTube page among the sources it retained.

Averaged across the four engines, YouTube would weigh 3.6% of citations: a number that describes no engine at all. On this corpus, one engine leans on videos in most of its sourced answers, one ranks the domain first while citing it sparingly, and two never credit it. Engine by engine is the only honest reading.

The part that gets read is the spoken track

For every source it retains, Perplexity's payload carries the passage the engine actually kept. That is what settles the question of which part of a video serves the answer.

All 205 video citations point to individual videos: never the homepage, never a channel page, never a results page. 121 distinct videos, 203 of them dated. And the retained passages are not descriptions: 129 of the 205 carry explicit second markers pointing into the video's timeline, and nearly all the rest reads as speech transcribed as heard, hesitations included. One retained passage recommends, in the transcript's own spelling, a "well ststructured website". Another, in French, advises creating content "optimisé pour lia", the machine's rendering of "l'IA". These are not editing accidents: they are proof of origin. The engine reads the transcript, machine generated when nothing better exists, and it reads deep: retained passages point as far as the sixteenth minute of a video.

That index is fresh. Half of the 205 entries carry a re-read date in July 2026, 71% since June. The pool of cited videos is young without being only new: 78% were published since the start of 2025, yet videos from 2020 still surface. And positions hold: 33 of the 118 identifiable videos come back on at least two distinct weekly surveys, a quarter of the pool re-cited week after week.

Gemini's payload exposes the domain rather than the page, but two things are visible through it. The search queries it runs are plain reformulations of the question. And its attribution ties entire blocks of the final answer to YouTube sources: lists of recommendations, selection criteria, practical advice, carried by what a video said. The infrastructure on Google's side is documented: Gemini grounds its answers through Google Search, and Google's video indexing reads what a video contains and detects its key moments, with caption indexing confirmed by the company for years.

Why these two engines, and not the others

The dividing line is access to the spoken track. A video's speech does not live in the HTML of its page: it is loaded by the player, through endpoints that youtube.com's robots.txt does not open to generic crawlers. An engine that reads the web page finds a title, a description, and little text worth quoting.

The two engines that quote YouTube are the two with a way in. Gemini belongs to the company that owns YouTube: its grounding runs on an index that treats videos as first-class content, transcripts and key moments included. Perplexity built its own video ingestion: the timestamped passages in our payloads are that index speaking. The two engines that do not quote YouTube encounter it as a web page: our archive shows Claude fetching a watch page and writing without it, ChatGPT never retaining one.

None of this is doctrine on any engine's part. It is the state of four retrieval pipelines, measured between July 10 and 24, 2026. Pipelines move without notice, which is a reason to measure rather than assume.

Titles open the door, transcripts do the talking

How does a video enter a source set in the first place? On this corpus: through its title. The videos our questions surfaced carry the question, nearly verbatim, as their title. Ask "What are the best places to visit in Africa right now?" and the source sets return "Top 10 Places To Visit in Africa" and "I Visited Every Country in Africa. Here's My Rankings". Ask "How can my business get recommended by AI assistants?" and they return "How to Get Your Business Recommended by EVERY AI Assistant". Even the two videos Claude went to fetch had restated the question in their titles. The title is what gets a video retrieved, on every engine that retrieves at all; what happens next is decided by the transcript.

Language follows the market. The French questions surfaced French-speaking videos first, English ones behind, and occasionally further afield: the matching crosses languages, but a video speaks to a market first in that market's language.

What this makes worth doing

Each property below aims at the mechanism measured above. Together they define what a video optimized for AI answers looks like.

  • Start from your own map. The engines weigh videos very differently, and the balance shifts by market and by question. Your own series, engine by engine, shows which sources serve the questions you care about: that is what decides how much a video deserves in your plan.
  • Give the video the title of the question. What enters the source set is a video whose title states the question your market asks, phrased the way it is asked. One question, one video, the title as the door.
  • Say the facts out loud. The transcript is the text the engines read. State your name distinctly: the automatic transcript writes what it hears, and what it writes is what gets indexed. Say the numbers, the names, the terms you want carried, in words spoken clearly.
  • Upload a reviewed transcript. Machine transcription is what the index falls back on. Captions you wrote yourself turn the indexed text into your own words, spelling included.
  • Chapter and describe. Key moments are detected and indexed. Chapters that restate the questions, and a description that answers them in clean text, give the detection something to find.
  • Speak the market's language. A video reaches a market first in that market's language; a second market deserves a second video more than a subtitle track alone.
  • Date the action, then read the effect. A published video is a move worth recording in your Logbook; the next surveys read against it. Your Atlas shows whether the sources behind your questions moved, and Epovest Tracking whether the answers themselves did.

A year ago, being quotable by AI engines mostly meant pages. On two of the four engines we measure, the most quoted source is now a place where content speaks rather than writes. The asymmetry itself is the lesson: visibility in AI answers is not one thing to win, it is four pipelines to measure, and our method starts there: ask the market's questions on a fixed cadence, archive the answers with their sources, and let the series say where to act.

Sources