SEO for AI means optimising websites so LLMs and AI search engines can retrieve, understand and cite them in generated answers. The foundations of traditional SEO still apply, but AI systems retrieve information rather than pages.
This guide covers how each major platform finds and selects sources, where Google and Microsoft officially disagree on tactics, and how to measure AI visibility using proprietary data from Roar’s property-sector tracking.
Introduction: Why are we talking about SEO for AI
Nearly 60% of Google searches now end without a click. And yet AI referrals to top websites grew 357% year on year, reaching 1.13 billion visits in June 2025. Both are true, which is exactly why SEO for AI has become the most confused topic in digital marketing.
Some agencies are selling llms.txt setup packages, while Google’s own guidance says the file is ignored. Microsoft recommends schema markup for Copilot visibility; Google’s documentation says no special schema is needed for its AI features. Half the industry is treating AI SEO as a brand new discipline. Google’s stance is that it’s still SEO.
Somebody is wrong, and it matters which one, because getting cited in an AI answer now sits alongside ranking as the visibility goal that determines whether a brand is even considered.
This guide covers what SEO for AI actually is, how retrieval works inside the major platforms, what to do (and ignore) on each one, and how to measure whether any of it is working.
What is SEO for AI?
SEO for AI is the practice of optimising a website so AI systems, including ChatGPT, Google’s AI Overviews, Perplexity and Copilot, can retrieve, interpret and cite its content in generated answers. It builds on the foundations of traditional SEO but targets citation and inclusion in AI responses, not ranking position alone.
Two other terms describe the same work from slightly different angles. Generative Engine Optimisation, or GEO, is the discipline of improving how often and how prominently AI systems cite a brand or website in generated answers. Answer Engine Optimisation, or AEO, focuses on structuring content so it appears as the direct answer in AI-generated responses, whether a click follows or not.
The industry has spent two years arguing about which label to use. Google settled the question in its official guidance for generative AI features: from its perspective, optimising for generative AI search is optimising for the search experience, and therefore still SEO.
That position is defensible, but it slightly understates the change. What has shifted is the outcome the work is measured against. Traditional SEO delivered rankings, which delivered clicks, which delivered leads. AI-mediated visibility increasingly delivers mentions and citations that influence a decision without a click ever happening. A brand referenced by an AI answer for a high-intent commercial query is being shortlisted, silently, in a decision process the marketing team may never see in analytics.
So the practical answer: call it SEO, GEO or AI SEO. The discipline is the same. The goalpost has moved.
How do LLMs and AI search engines find content?
AI search engines find content through retrieval-augmented generation, or RAG: when a query needs current information, the model generates multiple related searches, pulls relevant pages from a search index, and composes an answer from what it retrieves. Visibility depends on appearing across that cluster of generated searches, not just one keyword.
The mechanic that matters most here is query fan-out. Query fan-out is the set of related searches an AI model generates from a single prompt to improve the coverage and accuracy of its answer. When someone asks ChatGPT “what are the best SEO strategies for 2026”, the model does not run that phrase alone. It generates variations: “SEO trends 2026”, “how AI search is changing SEO”, “current best practices in search engine optimisation”, and pulls sources across all of them. A page that ranks well for the primary phrase but appears nowhere in the fan-out cluster gets ignored.
Grounding queries are the second half of the picture. Grounding queries are the actual searches an AI system runs while generating a response, as opposed to the fan-out variations it predicts it might need. Some tools now expose them. Microsoft Bing Webmaster Tools reports which grounding queries triggered a page as a Copilot source, which is one of the few ways to see, rather than guess, how AI retrieval treats a site. For a deeper walkthrough of the mechanic and how to observe it in practice, how AI answer engines like ChatGPT search for content covers the observation method step by step.
How is ranking in AI search different from ranking in Google?
Traditional search ranks whole pages in an ordered list. AI search parses pages into smaller passages, evaluates each one for relevance and authority, and assembles an answer from multiple sources. A page can rank first organically and still never be cited, because selection now happens at the passage level, not the page level.
That single change explains most of what feels different about AI SEO in practice. When a page is treated as a set of retrievable passages, three things follow.
The first is that the strongest 40 words of a page carry disproportionate weight, because those are the ones an AI system is most likely to extract. Content that buries the answer three paragraphs in loses to content that opens with it. The second is that citations and mentions become separate outcomes worth tracking on their own. A page can be quoted verbatim in an AI Overview with a linked citation, name-checked in the answer text without a link, or referenced silently as a source. Each has different commercial value. The third is that ranking well and being cited are no longer the same job. A page with mid-tier organic rankings can outperform the top result in AI citations if its passage structure is cleaner.
The scale of this is easy to underestimate. Nearly 60% of Google searches now end without a click, which means the assumption that a strong ranking produces a visit no longer holds for most searches. For commercial, decision-stage queries specifically, the shift is sharper: Google AI Overviews appeared on 87% of 500,000 commercial-intent prompts, and on 88.5% of prompts classified as decision-stage. The AI Overview is, in effect, the first commercial result on the page.
Traditional ranking has not stopped mattering. It has stopped being sufficient. Being on the page is now the entry requirement. Being cited in the answer is the outcome.
What actually works: SEO for AI best practices in 2026
The tactics that reliably improve AI search visibility in 2026 come down to four things: non-commodity content with first-hand insight, passage-level structure that puts direct answers under question headings, entity and brand signals earned across the web, and technical foundations that keep content crawlable by both search and AI crawlers.
Create content AI cannot get anywhere else
The strongest single lever is producing content that a language model cannot synthesise from what is already online. Google’s guidance calls this “non-commodity content” and uses a specific test: a generic listicle like “7 Tips for First-Time Homebuyers” restates common knowledge, while a first-hand piece explaining why a buyer waived a home inspection and saved money offers something the model cannot get anywhere else.
For property clients, that has meant leaning on real transaction data. A branch that publishes what actually happened when a valuation in a specific London postcode went to sealed bids last month is a source. A branch that publishes “10 tips for selling your home in spring” is not.
The rule of thumb is straightforward: if a competent AI could produce a page from the top ten Google results without visiting the site, that page has nothing to contribute to an AI answer. Its job now is to add something the model does not already know.
Structure every section to stand alone
AI systems do not read pages sequentially. They parse them into passages and evaluate each passage for whether it answers the query on its own. That mechanic changes how every section should be written.
The practical rule is that the first 30 to 60 words under any subheading should answer the question that subheading implies, without depending on earlier context. Pronouns that refer back to something two sections ago make a passage unquotable. Sentences that begin “as we saw above” or “building on this” do the same. The strongest AI-friendly writing is not deep-linked prose; it is a set of self-sufficient answers under clear question headings that a retrieval system can lift, quote and cite in isolation.
Question-format headings help for the same reason. “How does query fan-out work?” is easier for a retrieval system to match to a user prompt than “Understanding modern search behaviour”. The heading becomes retrievable in its own right.
Build entity and brand authority beyond your own site
An AI system deciding who to cite for a topic looks at more than a single site. It looks at whether the brand is referenced consistently across sources it already trusts. That is entity SEO. Entity SEO is the practice of establishing a brand, person or business as a recognised entity across the web through consistent naming, structured data and third-party references that reinforce who they are and what they do.
The mechanics are less mysterious than the label suggests. A brand mentioned in trade press, cited in industry reports, referenced in Reddit threads and quoted in journalist coverage becomes a stronger citation candidate than one that exists only on its own domain. Digital PR, expert commentary, unlinked brand mentions and E-E-A-T signals all feed the same outcome: the entity is recognised.
For clients considering working with an AI SEO agency, this is the piece most often underinvested. Content optimisation without brand building produces well-written pages that AI systems still overlook, because nothing outside the site vouches for them.
Keep your content technically accessible to AI crawlers
Nothing above works if the page cannot be crawled. Every AI system runs its own crawler and each one respects robots.txt independently:
- Google-Extended covers Gemini and Google’s generative AI features, separate from Googlebot
- GPTBot is used by ChatGPT for training data
- OAI-SearchBot handles retrieval for ChatGPT search
- PerplexityBot handles Perplexity’s index
- ClaudeBot is used by Anthropic’s Claude
Blocking any of these on a site that wants AI visibility is a self-inflicted wound, and one that a lot of large sites are still carrying accidentally. JavaScript rendering matters here too. Content that only appears after a client-side render is invisible to most AI crawlers, because they process the initial HTML response.
The base requirement is boring but non-negotiable: the page has to be indexed and eligible for a snippet in normal search. Everything else is optimisation on top of that.
What Google and Microsoft disagree on (and what to do about it)
Google’s official guidance says llms.txt files, content chunking, AI-specific rewriting and special schema markup are unnecessary for its AI features. Microsoft’s guidance for Bing and Copilot recommends structured, parseable content and schema markup as core practices. Both are correct for their own systems. Which approach a site should take depends on which platforms drive the audience.
The tension is not really disagreement in the strict sense. Each company is describing what its own systems reward. Google’s AI features draw on the existing Google Search index and ranking systems, so its position is that generative AI visibility is downstream of ordinary SEO. Microsoft’s Copilot parses content into passages that are evaluated for structure, and its documentation on optimising for AI search answers is more explicit about how that parsing works. Neither company covers ChatGPT, Perplexity or Claude directly, though ChatGPT and Copilot both retrieve through Bing.
Where this matters most is llms.txt. llms.txt is a proposed plain-text file that lists a site’s key content for large language models, similar in concept to robots.txt. Google has confirmed it ignores llms.txt entirely, along with the wider “AI-specific” markup ecosystem being sold as necessary. Microsoft has not commented on it. There is no evidence that any major AI platform meaningfully weights it in retrieval.
The full verdict, tactic by tactic:
| Tactic | Google’s position | Microsoft’s position | Our verdict |
|---|---|---|---|
| llms.txt files | Ignored, not used | Not addressed | Low priority. Harmless if trivial to add, but no evidence of impact on any major platform |
| Content chunking for AI | Not required; systems handle multi-topic pages | Recommends modular, parseable content | Write for the reader first. Clear question-format H2s and self-contained passages help Bing and Copilot without hurting Google |
| Schema markup | Not required for AI features; useful for rich results | Recommends structured data | Keep it. Semantic clarity supports every platform and rich results eligibility remains valuable |
| AI-specific content rewriting | Unnecessary; systems understand synonyms | Not explicitly addressed | Skip it. Optimising for what an LLM is presumed to expect is a distraction |
| Inauthentic brand mentions | Flagged as unhelpful and filtered | Not addressed | Avoid everywhere. Authentic mentions build entity strength; manufactured ones get suppressed |
The practical calculation for most UK businesses is straightforward. Google still drives the majority of commercial demand, so Google’s guidance carries most weight on tactics that would cost time or money to implement. Where Microsoft’s guidance adds structural discipline that also serves human readers, follow it. Where it recommends effort that Google explicitly ignores, deprioritise it.
The one point both companies agree on, quietly, is that authentic content is worth more than any technical layer bolted on top. That agreement is the one worth building strategy around.
How to optimise for each AI platform
Each AI platform selects sources differently. Google’s AI Overviews and AI Mode draw on Google’s index and its core ranking systems. ChatGPT and Copilot retrieve through Bing. Perplexity runs its own crawler with a bias toward recently updated, well-structured pages. Optimising across all four means being visible in both major search indexes and earning citations in the sources each system trusts.
Google AI Overviews and AI Mode
Retrieval source: Google’s index, filtered by its core ranking and quality systems. Selection behaviour: AI Overviews favour pages that already rank well for the underlying query and cite recent, authoritative sources with clear passage structure. AI Mode extends the same logic across multi-turn conversations, so a page can be cited across a conversation rather than just a single query.
The two most useful actions for a Google-focused site are:
- Confirm the site is enabled for generative AI features in Search Console, and check the Generative AI performance report for citation appearances
- Structure content so the first passage under each subheading answers the implied question in isolation, which is what an AI Overview lifts
Everything else, including llms.txt and AI-specific markup, Google has said explicitly is not needed.
ChatGPT
Retrieval source: Bing, via OAI-SearchBot. Selection behaviour: ChatGPT search runs query fan-out and pulls from multiple Bing results per prompt, weighting recency and semantic relevance. Historic training data still influences the model’s baseline familiarity with a brand, but live citations come from what Bing surfaces at the moment of the query.
Two actions carry most of the return:
Register and verify the site in Bing Webmaster Tools. Bing indexation is the entry requirement for ChatGPT search visibility, and Bing’s index is materially different from Google’s for many UK sites
Publish content that answers commercial-intent questions clearly and openly. ChatGPT is measurably more likely to cite pages that answer the question in the first passage rather than gating the answer behind lead capture or long preambles
Perplexity
Retrieval source: PerplexityBot, its own crawler. Selection behaviour: Perplexity favours recently updated content, cites its sources visibly in the interface, and shows a stronger preference for original data, first-hand experience and structured content than any other major platform in 2026. It also indexes Reddit and other community sources aggressively.
The single most useful action is publishing citable, dated original content: proprietary data, first-hand observations, methodology documentation. Perplexity users click through to sources more often than users of any other AI platform, which makes it disproportionately valuable for lead generation on informational queries.
Microsoft CoPilot
Retrieval source: Bing, same index as ChatGPT search, plus Microsoft’s ecosystem signals from Merchant Center and Business Profiles for commercial and local queries. Selection behaviour: Copilot places heavier weight on structured content, schema markup and passage-level parsing than Google does, in line with Microsoft’s published guidance.
Two useful actions:
- Verify Bing Webmaster Tools access and monitor the grounding queries report, which is the only publicly available data on which searches actually trigger a page as an AI source
- Maintain schema markup and semantic HTML, which Copilot rewards even where Google says it is not required. This is the clearest case where following both companies’ guidance means following Microsoft’s
Gemini and Claude are worth a brief note. Gemini shares Google’s index, so most Google AI Overview work covers it. Claude retrieves through Brave Search when live retrieval is enabled, which means Brave indexation matters for Claude visibility, though Claude’s overall query volume remains smaller than the four platforms above.
How do you measure AI search visibility?
AI search visibility is measured by tracking how often a brand is cited or mentioned across AI-generated answers for its target queries, not by rankings or sessions. The tools available in 2026 are Google Search Console’s Generative AI performance report, Bing Webmaster Tools’ grounding data, and dedicated AI visibility tracking that scores citations independently of clicks.
Standard analytics miss most of what matters. Google Analytics reports the visits a page received, not the answers a brand was cited in without a click. Search Console tells you a page appeared in results, not whether it was quoted in an AI Overview. For any brand competing on commercial queries, that measurement gap is now the single biggest blind spot in the marketing stack.
What Google Search Console can and cannot tell you
Google’s Generative AI performance report shows which pages surfaced in AI features on Google Search and Discover, alongside impressions and clicks from those appearances. It is the only first-party data available for AI Overview and AI Mode visibility, so a site not verified in Search Console is measuring nothing on the Google side.
What it does not show is which specific query triggered a citation, whether a page was quoted verbatim or referenced silently, or how a brand compares to its competitors in the same AI answers. Bing Webmaster Tools fills part of that gap for Copilot and ChatGPT search by exposing grounding queries, the actual searches that pulled a page as a source. Between the two platforms, most first-party AI visibility data available in 2026 comes from Google and Microsoft’s own consoles.
Tracking citations and mentions: the AI Visibility Score
Where first-party data ends, dedicated visibility tracking begins. The core problem is that an AI answer can reference a brand in three different ways, and each has different commercial weight. A brand can be cited as a linked source, named in the answer text without a link, or both.
Roar tracks this using an AI Visibility Score with three weightings:
- Both, where a brand is named in the answer text and cited as a linked source: scored 1.0
- Source only, where a brand appears as a linked citation without being named in the text: scored 0.5
- Text only, where a brand is named without a link: scored 0.5
The full formula is straightforward: AIS = (1.0 × Both) + (0.5 × Source Only) + (0.5 × Text Only).
Running that score across a defined set of target queries every month produces a visibility measure that operates independently of clicks. It has become the metric most Roar property clients now report on internally, because it captures the commercial reality that a shortlist mention with no click is still a shortlist mention.
A note on the data behind that method. Roar has been running AI Overview visibility tracking across 4,000+ UK locations in the property sector for the last twelve months, alongside four years of organic coverage data in the same market. That is one sector, not a universal benchmark, and the piece of data that stands out most from it is volatility. AI Overview citations in property change enough month to month that a single-month reading is misleading. Peec AI’s broader dataset supports the same pattern: AI Overview presence varies from 64.6% for two-word queries to 88.5% for decision-stage prompts. Query mix moves the number as much as anything a brand does.
The practical implication is that AI visibility should be reviewed quarterly, on a defined query set, using a scored method that treats mentions and citations as separate outcomes. Anything less granular hides most of what is actually happening.
Common SEO for AI mistakes to avoid
The most common SEO for AI mistakes in 2026 are publishing generic AI-written content with no original insight, buying llms.txt or chunking services for Google visibility, chasing inauthentic brand mentions, treating AI SEO as separate from technical SEO, and measuring success with traffic alone while citations go untracked.
Each one is worth being specific about.
Publishing generic AI-written content is the mistake that shows up most often, and the one AI systems are most efficient at ignoring. If a page can be produced by prompting an LLM with the top ten Google results, the same LLM will have nothing to gain by citing it. Volume is not the problem. Sameness is.
Buying llms.txt or AI-specific markup services for Google visibility is the second. Google has confirmed it ignores llms.txt along with the wider “special markup” ecosystem some agencies are selling. There is no evidence that any major AI platform meaningfully weights it. Paying for it as a Google visibility play is paying for nothing.
Chasing inauthentic brand mentions is the third, and it is one of the few areas where Google has been explicit about active filtering. Manufactured mentions across low-quality blogs, comment sections and forum spam are recognised by ranking systems and suppressed. Real mentions, earned in publications and communities that already carry weight, are the ones that build entity strength.
Treating AI SEO as separate from technical SEO is the fourth. A page that AI cannot crawl, index or render is invisible regardless of how well it is written. The base requirements are the same as they have always been. If Google or Bing cannot fetch the page, no AI system will cite it.
Measuring success with traffic alone is the mistake most likely to catch out marketing teams that have been strong on traditional SEO. Zero-click AI answers still influence commercial decisions, and a brand that only tracks sessions is invisible to itself in the exact channel that increasingly decides consideration. This is the case for benchmarking AI visibility separately from organic performance, using a scored method that treats citations and mentions as first-class outcomes.
There is a common thread across all five. Each mistake treats AI SEO as either a new set of tricks or a lighter version of the old rules. It is neither. It is the same discipline, executed to a higher standard, measured against different outcomes.
Conclusion:
The tension that runs through this whole guide is that AI search rewards the same things good SEO always has, executed to a higher standard and measured differently. There is no separate playbook. There are just fewer places to hide when a page is thin, a technical foundation is weak, or a brand has no authority beyond its own domain.
Three things are worth taking away.
The disagreement between Google and Microsoft is smaller than it looks. Write for the reader, structure passages so they answer questions in isolation, keep schema for the reasons it has always been useful, and skip the tactics being sold as AI-specific shortcuts.
Second, visibility now has to be measured against citations and mentions, not just clicks. Any brand competing on commercial queries in 2026 without a scored view of its AI visibility is measuring itself in the wrong currency. And the platform mix matters. Being visible in both Google and Bing is now the entry requirement for AI answers across the ecosystem, not an optional extra for Bing loyalists.
If any of the above is unclear where your own site is concerned, learn about GEO/ AI SEO solutions or get in touch. The gap between brands that adapt to this and brands that do not is widening quickly, and the ones that move first are the ones that stay in the answer.
Frequently asked questions
Is SEO dead now that AI search exists?
No. AI search systems are built on top of search indexes and core ranking systems, so traditional SEO remains the entry requirement. What has changed is that a strong ranking no longer guarantees visibility on its own, because content also has to be selected and cited at the passage level inside AI-generated answers.
What is the difference between SEO, AEO and GEO?
All three describe optimising for search visibility. AEO, or Answer Engine Optimisation, focuses on structuring content so it appears as the direct answer in AI responses. GEO, or Generative Engine Optimisation, focuses on being cited by generative AI systems in the answers they produce. Google’s official position is that optimising for its generative AI features is still SEO, and in practice the three disciplines overlap almost entirely.
Do I need an llms.txt file?
No, not for Google. Google has confirmed it ignores llms.txt along with other AI-specific markup files. The file is harmless to add and some smaller systems may read it in future, but it belongs at the bottom of any priority list, well below content quality, entity strength and technical accessibility.
How do I get my business mentioned in ChatGPT?
ChatGPT retrieves live information through Bing, so verified Bing indexation is the foundation. Beyond that, earn authentic mentions and citations in the sources ChatGPT trusts for your topic, and structure content so individual passages answer commercial-intent questions completely rather than hiding the answer behind a preamble.
How often should AI visibility be reviewed?
Quarterly, on a defined query set. AI-generated answers change far more frequently than organic rankings, and query mix moves the number as much as any optimisation activity. Peec AI’s data shows AI Overview presence varies from 64.6% for two-word queries to 88.5% for decision-stage prompts, which means a single-month reading on a narrow query set can mislead badly.
Does schema markup help with AI search?
It depends on the platform. Google states no special schema is needed for its AI features. Microsoft recommends structured data for Copilot visibility. Schema remains worth maintaining for rich results eligibility and semantic clarity across every platform, but it should not be treated as a standalone AI visibility lever on Google.