What LLM SEO means
LLM SEO is the practice of optimizing how large language models retrieve and recall your brand. Every AI answer that names a brand arrives through one of two mechanical paths: the model recalls the brand from its training data, or it retrieves fresh pages through a live search and cites them. LLM SEO works on both paths — shaping what models remember about your brand and what they find when they look. For the broader strategy discipline built on top of these mechanics, see our guide to what GEO is; this page stays underneath the strategy, on how models actually find and remember brands.
Recall vs. retrieval: the two paths to an answer
Recall is parametric: the model answers from what it absorbed during training, with no live lookup and no citation trail. What it 'knows' about your brand was fixed when the model was trained, and it changes only when a newer model trains on a corpus where your brand appears — which makes recall slow to influence but persistent once earned. Retrieval is the opposite profile: the engine runs a search, reads the top results, and synthesizes an answer with citations attached. It responds to fresh, well-ranked, cleanly extractable content far faster than recall can change, but the result varies per query and per session. A single conversation typically blends both paths, so a mature program invests in both.
- Recall: sourced from training corpora; slow to change; persists across every conversation; leaves no citation trail
- Retrieval: sourced from live search; responsive to new content; visible through citations; varies by query and session
- Recall failures (wrong category, stale facts) persist until models retrain — correct the public record early
- Retrieval failures (never cited) are diagnosable — the citations show exactly which sources beat you
Citation sources differ by engine
There is no universal source mix. Muck Rack's May 2026 'What Is AI Reading?' analysis of more than 25 million links across ChatGPT, Claude, and Gemini in 17 industries found each engine leans on a different backbone: Wikipedia was the top cited domain for ChatGPT, PubMed Central for Claude, and Reddit for Gemini. The same study found earned media made up 84% of links while paid or advertorial content contributed just 0.3%, with journalism supplying 27% overall — and press releases appeared 3.5 times more often for trend prompts than for best-of prompts. Profound's separate 27-million-citation analysis adds a per-engine social split: social sources represented 19% of Perplexity citations but only 5% of ChatGPT's and 7% of Google AI Overviews', and social citations rose from 5.4% on category prompts to 15% on brand prompts. Owned sources supplied only 4.3% of citations for category prompts, so your own site cannot carry visibility alone — citations come from a portfolio of editorial, community, institutional, and review sources, weighted differently per engine.
- ChatGPT: Wikipedia is the top cited domain (Muck Rack, May 2026)
- Claude: PubMed Central tops the citation list
- Gemini: Reddit is the most cited domain
- Earned media: 84% of analyzed links; paid or advertorial content: 0.3%
- Owned sites: only 4.3% of citations on category prompts (Profound)
How retrieval relates to search rankings
The retrieval path routes through traditional search rankings — but unevenly by engine. Semrush studied 5,000 queries and 150,000 citations and found Perplexity cited a domain from Google's top ten in 91% of cases and the exact URL in 82%; for Google AI Overviews the figures were 86% and 67%. A separate Ahrefs test of 3,311 head terms found much lower ChatGPT overlap — 31.8% at the domain level and just 10% at the exact-URL level — while its Perplexity overlap was 80.58% and 65.07%. The practical read: Perplexity and AI Overviews citations overlap heavily with Google's top results, while ChatGPT's citations are far less correlated with Google rankings. These are overlap figures, not proof of mechanism — the studies frame traditional search visibility and passage-level relevance as complementary, so the defensible play is a rank-worthy page with a self-contained passage that precisely answers the target question. Semrush also found commercial AI responses ran roughly twice as long as informational ones, so pages targeting commercial retrieval need the depth — comparisons, trade-offs, evidence — to feed a longer synthesized answer.
Citations decay as conversations deepen
Being cited is largely a first-turn game. Profound analyzed roughly 730,000 cited US-English conversations from October through December 2025 and found citation incidence fell from 12.6% at turn one to 4.5% at turn ten and 3% at turn twenty; a cited conversation averaged about six citations. The study reports the decline without establishing why it happens — one plausible interpretation is that as a session deepens, the model increasingly reasons over what it has already retrieved and recalled rather than searching again, but that mechanism is inference, not measurement. The implication for content is direct either way: build pages that resolve the opening research questions — what is, best for, how much, X versus Y — in the first extractable passage, because that is where the citation opportunity is concentrated.
Entity consistency: what the model remembers
The recall path does not store your pages; it stores associations — your brand name linked to a category, a description, and related entities, learned from every crawled mention. If the public record describes your brand inconsistently — different category labels on different profiles, conflicting boilerplate, aliases that never resolve to one name — the model's recalled representation is diluted or simply wrong, and no single page fixes it. The work is unglamorous: settle one canonical description, use it verbatim across your site, directories, review profiles, and press materials, keep naming consistent so aliases collapse to one entity, and ensure the description is corroborated by independently crawled sources rather than existing only on your own domain.
- Write one canonical one-sentence brand description and reuse it verbatim everywhere
- Standardize the brand name so variants and aliases resolve to a single entity
- Keep category labels identical across directories, review sites, and press boilerplate
- Get the canonical description corroborated on independently crawled third-party sources
Verify citations — and separate vendor theories from confirmed guidance
Citations are evidence, not proof. A 2026 preprint (arXiv 2605.14021) that studied 55,393 queries across 19 categories over 40 days reported that about 30% of cited domains were not on the first search-results page and that 11% of 98,020 atomic claims were unsupported by their attached citations — provisional figures, since the paper is under review, but a good reason to audit what citations actually say rather than only counting mentions. On the Google side, official guidance is unusually clear: the same foundational SEO practices apply to AI features, no special AI text files or special schema are required, inauthentic mentions are warned against, and third-party tools do not have access to Google's internal search data. Treat llms.txt and similar conventions as unproven vendor hypotheses, not requirements. Google also added Search Generative AI performance reports to Search Console in June 2026 for a subset of sites, covering AI Overviews and AI Mode by page, country, device, and date — a free, first-party way to measure the retrieval path on Google surfaces.