AI discovery
How AI assistants choose which sources to cite
When AI assistants search the web, they run one or more searches based on the question, retrieve pages they are allowed to crawl, and cite the pages they used to build the answer.
Google describes this as query fan-out. Ranking in the underlying search, crawl access, and whether a page clearly answers part of the question all matter. The exact selection logic is not public, and results vary between runs.
What is documented
The AI companies publish only part of how this works. What they do publish is useful.
- Google. Its documentation for AI Overviews and AI Mode says these features may use a "query fan-out" technique, running multiple related searches across subtopics and data sources, and that this can surface a wider set of links than classic search. Google also says that pages must be indexed and eligible to show a snippet in search to be shown as a supporting link.
- OpenAI. OAI-SearchBot crawls pages for ChatGPT's search features. Sites that block it are not shown in ChatGPT search answers. ChatGPT-User makes requests when a user asks it to visit a page.
- Perplexity. PerplexityBot indexes sites to surface and link them in Perplexity's answers, and Perplexity-User fetches pages on demand for a user's question.
A simple model of the process
- The assistant interprets the question and decides whether to search.
- It generates one or more search queries, often more specific than the user's wording.
- It retrieves candidate pages from a search index.
- It reads the parts of those pages that relate to the question.
- It writes an answer and cites the pages it relied on.
This is a simplification. It helps explain why a page that ranks for a narrow sub-question can be cited even if it does not rank for the original question.
What tends to matter
From the documentation and published studies, these factors appear repeatedly:
- Crawl access. If the crawler is blocked, the page cannot be used.
- Search visibility. Pages that rank for the sub-queries are more likely to be retrieved. The relationship is not one to one; analyses report that AI citations increasingly come from outside Google's top 10 results.
- A clear answer to a specific question. A page that answers one question directly is easier to use than a page that mentions it in passing.
- Source type. Studies by tracking vendors find assistants cite third-party sources heavily: reviews, forums, video, news and professional profiles. Digiday, using Meltwater data, reported YouTube as the most-cited platform across eight AI products in August 2026.
What is not known
- The weighting of each factor
- How much the model's training data influences which sources it trusts
- How often each system's retrieval changes
Be wary of anyone who claims precise knowledge of the algorithm. The published research is mostly from vendors that sell tracking tools, and its findings shift from month to month.
Engines differ
ChatGPT, Perplexity and Google draw on different indexes and show citations differently. Perplexity shows sources prominently on every answer. Google's AI features sit inside search results. ChatGPT cites when it searches. A company can be well cited in one and missing in another, which is why testing each separately matters. See how to test whether AI assistants know your company.
What to do with this
The practical actions are the same ones that help search in general: allow the relevant crawlers, publish pages that answer specific buyer questions, keep your company information consistent, and be present in the third-party sources your buyers and the assistants trust.
Sources
- AI features and your website, Google Search Central, 2025-12-10.
- Overview of OpenAI crawlers, OpenAI.
- Perplexity crawlers, Perplexity.
- In graphic detail: LLMs keep citing YouTube in search results, Digiday, 2026-09-23. Uses Meltwater citation data.
- GEO and AEO: what the evidence supports, Papercrane. Summarises Ahrefs llms.txt data and other studies. Secondary source.