Epovest
← All studies

Study

What AI assistants actually cite: the site profiles behind the answers, engine by engine

By Simon Vasconcelos Lee

When an AI assistant answers a market's question, whose pages is it standing on? The names that come to mind are the big ones: legacy media, the household platforms. If that were true, the door would already be closed for everyone else. So we measured it, on our own panels: 923 answers archived with every source they cite, across four engines and five markets, in July 2026.

Two results carry this study. The cited web is far more open than the "always the same big names" assumption suggests. And each engine trusts a visibly different web, so every number below is split by engine, never averaged.

The measurement

Between July 10 and July 25, 2026, our trackers put 18 fixed question panels to ChatGPT, Claude, Gemini and Perplexity through their official APIs, the same way they run for any account. The panels cover five markets: travel destinations (648 answers), SaaS and AI visibility (208), web services in French (27), currency data (16), and brand identity questions (24). Every answer is archived with its cited sources: 6,991 citations in total, pointing at 2,049 unique root domains.

Travel and SaaS carry the statistical weight. The three small panels are read as indications, not laws, and are flagged as such below.

Finding 1: no oligopoly holds the answers today

For each engine, we summed the citations taken by its ten most-cited domains:

Engine Travel questions SaaS questions
ChatGPT 43% 32%
Perplexity 28% 23%
Claude 25% 27%
Gemini 12% 16%

Even at the most concentrated end (ChatGPT on travel), the top ten domains leave 57% of citations to everyone else. At the open end, Gemini's ten most-cited travel domains carry 12% of its citations, and it took 746 distinct domains to carry the rest. The set of pages AI answers stand on is a long tail, and modest sites are inside it.

Finding 2: each engine trusts a different web

ChatGPT leans institutional. On travel, its three most-cited sources are Lonely Planet (129 citations), National Geographic (102) and UNESCO (51), with government domains close behind. Its bar is reputation built elsewhere, which makes it the slowest set to enter.

Gemini spreads the widest. Specialist operators, personal blogs, YouTube and Reddit share its top, and no source dominates. A focused site with one strong page on the exact question can enter its cited set: this is where new entrants surface first.

Perplexity reads platforms first. On travel questions, YouTube is its most-cited source (152 citations) and Reddit its second (138). The same pattern holds on SaaS questions, where LinkedIn also enters its top sources. As we measured in our YouTube study, what the engine actually reads there is the transcript.

Claude cites the least and reads the most. 466 citations across 233 answers, but on SaaS questions alone it read 342 pages without citing them, a figure only Claude exposes. When it does cite, the answers stand on independent consultants' blogs and agencies' pages: expertise signed by a person or a practice, rather than mass media.

On SaaS questions, one more fact matters to any vendor: tool makers' own pages are cited by every engine, alongside third parties. Answering your market's question on your own page is not wasted ink.

Finding 3: when the engine does not search, its memory answers

Not every answer triggers a web search. Measured on this corpus: Perplexity searched in 100% of its answers, Gemini in 96%, ChatGPT in 73%, Claude in 42%.

Claude's figure hides the most useful detail: it searches by subject. On SaaS and AI visibility questions it searched 79% of the time; on travel it searched 27% of the time and answered from memory otherwise. Where the engine believes it already knows, the training corpus answers, and that corpus was frozen months ago.

So there are two different games in one visibility question. Live retrieval rewards the page that answers the exact question asked, and can be won this week. The frozen corpus rewards what existed, stably and repeatedly, when the corpus was built, and only updates when the next model ships. Publishing early and keeping URLs and facts stable is how a brand plays both.

Finding 4: trust questions anchor on modest third parties

Our brand identity panel is small (24 answers, indicative), but its pattern is sharp and matches what we measured in our backlinks study: when engines answer "can I trust this company", the citations go to review platforms (Trustpilot first), LinkedIn, and third-party registries and verification sites, not to the press. The third-party record is what carries the verdict, and those records are pages any company can have.

What a brand does with this

  1. Cover your market's questions on your own pages. Vendor pages are cited when they answer the question asked, and they are the one set of pages you fully control.
  2. Exist where each engine reads. A well-titled video whose transcript answers the question, for Gemini and Perplexity. A clean third-party record, for trust questions. Expert-signed pages, for Claude.
  3. Measure per engine, never on average. Every number in this study splits four ways. A strategy tuned on the average optimizes for an engine that does not exist.

Method

923 answers collected between 2026-07-10 and 2026-07-25 by Epovest trackers through the engines' official APIs, on fixed panels phrased as customers ask them. Every answer is archived with its cited sources; domains are normalized to their registrable root (2,049 unique roots). Cell sizes vary by market and engine and are given in the text. Figures may be reused with attribution.