Your website now has two audiences. One is a person. The other is a machine fetching your page on that person’s behalf, summarizing it, and deciding whether to name you.
The second audience is growing fast, and it is far less forgiving than the first. A human will squint at a slow page, scroll past a wall of text, and figure out what you sell. A retrieval system will not. It takes what it can parse, grounds an answer in that, and moves on.
The good news is that most of the work is not exotic. In May 2026, Google published its first official guidance on optimizing for generative AI features in Search, and the headline is deflating for anyone selling a separate AI playbook. Asked whether SEO still matters for generative AI search, Google’s answer is <a href=”https://developers.google.com/search/docs/fundamentals/ai-optimization-guide”>”In short, yes!”</a> AI Overviews and AI Mode run on the same index and the same ranking systems as the blue links.
So this is not a guide to tricking a model. It is a guide to building a site that is legible, fast, factual, and open to the systems that matter. Here is the order of operations.
1. Start with eligibility, not cleverness
Nothing else matters if your pages cannot be crawled and indexed.
Google’s AI features documentation is explicit that there are no extra technical requirements for AI Overviews or AI Mode. A page needs to be indexed and eligible to appear with a snippet, which means meeting the standard Search technical requirements. That is the whole bar.
What to check:
- robots.txt is not quietly blocking you. Review the robots.txt basics and confirm nothing important is disallowed.
- Your CDN or WAF is not blocking crawlers either. This is a common and invisible failure. Bot-mitigation rules added by a security vendor can block legitimate crawlers without anyone in marketing knowing.
- You have not opted out. Sites must be included in generative AI features in Search Console to be eligible for display in them.
- Duplicate content is under control. Duplicates waste crawl budget on URLs nobody cares about. Google’s SEO Starter Guide covers the basics of reducing it.
- JavaScript is not hiding your content. Google can process content inside JavaScript when it is not blocked, but JS-heavy frameworks make everything harder. Follow the JavaScript SEO basics.
That last point deserves emphasis. Many AI crawlers outside Google do not execute JavaScript at all. If your key product copy only exists after hydration, some systems see an empty shell. Server-side rendering or static generation is now a visibility decision, not just a performance one.
2. Put the answer in text, near the top, in plain HTML
Retrieval systems work with text. Anything you lock inside an image, a PDF, a video with no transcript, or a slide widget is effectively invisible.
Practical rules:
- Lead with a direct, self-contained answer. Not a brand narrative, not a scene-setting paragraph. A model extracting a claim wants a sentence that stands on its own without the three paragraphs above it.
- Use real headings and real paragraphs. Google advises organizing content with clear sections and headings because it helps readers navigate, and that same structure helps machines segment your page.
- Use semantic HTML where it is easy. Google notes that perfect semantic markup is not required, but it helps screen readers and agents parse your page. Treat it as an accessibility win with a side benefit.
- Do not chunk your content into confetti. Google explicitly says there is no requirement to break content into tiny pieces for AI, and no ideal page length. This is one of the more persistent pieces of bad advice circulating right now.
3. Write the thing only you can write
This is the part Google names as the biggest lever, above any technical change.
The guidance draws a line between commodity content and non-commodity content. A generic listicle assembled from common knowledge is commodity content, and a language model can produce it instantly. First-hand experience, original data, a specific decision you made and what it cost, an expert explaining a process they actually run: that is the material a synthesis engine has to source from somewhere.
If your page can be reproduced by an LLM from its training data, there is no reason for an LLM to cite your page.
Practical version:
- Interview your own subject-matter experts before drafting anything technical.
- Publish the questions customers actually ask, with real answers, including the unflattering ones.
- Put named authors with credentials on pages where expertise matters.
- Show your work. Numbers from your own operations, timelines from real projects, and specifics from real installations cannot be synthesized from elsewhere.
Google’s fuller framing is in Creating helpful, reliable, people-first content.
One caution worth naming: mass-producing pages to cover every possible query variation is treated as scaled content abuse. The “cover every fan-out query” tactic making the rounds is a spam policy violation, not a strategy.
4. Use structured data, but understand what it does
Schema markup is one of the most oversold tactics in AI search discourse. Google’s position is that structured data is not required for generative AI search and there is no special schema for it. It remains worth doing because it makes you eligible for rich results in regular Search.
The genuinely useful version is entity clarity. Mark up who you are, what you sell, where you operate, and what things cost, so there is an unambiguous machine-readable statement of fact on your own domain. When a model has to choose between guessing and reading, give it something to read.
Start here:
- Intro to structured data
- Organization schema
- Local business schema
- Product structured data
- Validate with the Rich Results Test
One hard rule: your structured data must match the visible text on the page. Mismatches are a policy problem, not a clever shortcut.
5. Make the page pleasant to land on
Page experience is not a magic AI signal, but it affects whether the person who does click stays. Google’s guidance asks for sites that render well across devices, reduce latency, and make the main content easy to distinguish from everything else around it.
References: Understanding page experience, Core Web Vitals, and PageSpeed Insights for measurement.
This matters more than it used to, because there are fewer clicks to work with. Pew Research Center’s analysis of roughly 69,000 Google searches across 900 U.S. adults found that when an AI summary appeared, users clicked a traditional result 8% of the time, compared with 15% when no summary appeared. Clicks on the sources cited inside the summary happened in about 1% of visits. The full report is here. When the clicks get scarcer, wasting one on a slow page is more expensive.
6. Do not leave images, video, and local data on the table
Generative AI experiences surface more than web page links. Images and video are additional slots your site can occupy.
- Follow the image SEO best practices and the video SEO documentation.
- Publish transcripts for video. A transcript is a text document that happens to have a video attached, and it is indexable in ways the video alone is not.
- For local businesses, keep Google Business Profile accurate, and see adding and managing business details.
- For ecommerce, product data through Merchant Center feeds into both AI responses and standard results.
7. Decide what non-Google AI systems are allowed to do
Google is one surface. ChatGPT, Claude, Perplexity, and Copilot fetch pages with their own user agents and their own rules, and your robots.txt is where you make that call.
The distinction most teams miss is between training crawlers and live retrieval crawlers. Blocking a training crawler keeps your content out of a future model’s weights. Blocking a retrieval crawler keeps you out of today’s answer. Those are very different business decisions and they are often made accidentally in a single copy-pasted robots.txt block.
Roughly, the agents fall into these groups:
- OpenAI: GPTBot (training), OAI-SearchBot (ChatGPT search), ChatGPT-User (user-initiated fetch)
- Anthropic: ClaudeBot, Claude-SearchBot, Claude-User
- Perplexity: PerplexityBot, Perplexity-User
- Google: Google-Extended (controls some non-Search uses, separate from Googlebot)
Check each provider’s current published documentation before writing rules, because these names and behaviors change. Then verify what your robots.txt actually does, rather than what you think it does.
What about llms.txt?
llms.txt is a proposed Markdown file at your site root, introduced by Jeremy Howard of Answer.AI in September 2024, that gives AI systems a curated map of your most important pages. It contains no access directives. It is a routing hint, not a permission system.
Be clear-eyed about it. Google states plainly that you do not need llms.txt or any special machine-readable AI file to appear in its generative AI features, and that Google Search ignores these files entirely. It neither helps nor hurts there. No major AI provider has confirmed it as a ranking or citation input, and independent crawl-log analyses suggest AI bots rarely request the file today.
The honest recommendation: if you run a documentation site or a large knowledge base, it is a cheap experiment with a plausible payoff for agent navigation. For a typical marketing site, it is near the bottom of the priority list, well below crawlability, content quality, and page speed. Anyone selling it as the key to AI visibility is selling something.
8. Think about agents, at least a little
AI agents now visit sites to complete tasks, not just to read. They analyze visual renderings, inspect the DOM, and interpret the accessibility tree. If your checkout, booking flow, or spec comparison only works through custom JavaScript widgets with no accessible labels, an agent cannot use it.
This is early, and it is not urgent for most sites. If you want to get ahead of it, Google points to the agent-friendly website best practices guide on web.dev, and to emerging protocols like the Universal Commerce Protocol.
The practical takeaway for now: accessible markup is agent-ready markup. The work you do for screen readers pays off twice.
9. Measure it without fooling yourself
Google added a Generative AI performance report in Search Console, which shows how content performs specifically in generative AI features rather than leaving you to infer it from total traffic.
Alongside that, build a manual prompt panel. Take 30 to 50 real buyer questions, run them monthly across ChatGPT, Gemini, Perplexity, and Google AI Mode, and log whether you appear, how you are described, and who appears next to you. Accuracy is its own metric here. Being described incorrectly is a different failure than being absent, and it needs a different fix.
A warning worth repeating, because Google says it directly: no third-party tool has access to Google’s internal ranking or AI systems. Any vendor claiming “internal” metrics or guaranteed AI rankings is making it up. Google’s guidance on evaluating third-party SEO advice is a useful filter.
Where the advice genuinely conflicts
Two of the loudest claims in AI search marketing are contested, and you should know that before you build a budget around either.
Brand mentions. Vendor correlation studies, most prominently Ahrefs’ analysis of 75,000 brands (link to the current version on the Ahrefs blog before publishing, since they have updated it more than once), report that branded web mentions correlate with AI visibility far more strongly than backlinks. Google’s own guidance pushes back, warning against chasing inauthentic mentions and noting that its core ranking systems and spam systems both feed the AI features. Both can be partly true: authentic mentions probably reflect real brand authority that the systems pick up, while manufactured ones are what Google is warning about. What is clear is that a campaign to buy mentions is a bad idea, and that being genuinely well-covered is not.
Who is publishing the research. Most AI visibility studies come from companies selling AI visibility tools. Ahrefs, Semrush, Similarweb, and SE Ranking all have a product in this category. Their data is useful and it is also self-interested. Pew and Google’s own documentation are the sources to weight most heavily, for opposite reasons: one has no commercial stake, and the other is describing its own system.
The short version
If you are building or rebuilding a site this year, the sequence is:
- Confirm crawlability and indexing, including your CDN rules and JavaScript rendering.
- Get the real answer into plain text, high on the page, under a clear heading.
- Publish things only you can publish, with named experts attached.
- Add structured data that matches the page, for entity clarity and rich results.
- Make the page fast and comfortable to land on.
- Cover images, video with transcripts, and local or product data.
- Make a deliberate decision about non-Google AI crawlers in robots.txt.
- Keep markup accessible so agents can navigate it.
- Measure in Search Console and with a manual prompt panel, and distrust vendors claiming internal metrics.
Notice that eight of those nine are things a good web team was already supposed to be doing. That is the real finding. AI search did not replace the fundamentals. It raised the penalty for skipping them.


