How to Optimize Your Website for Answer Engines

Portrait of Louise Linehan

By Louise Linehan

Content Marketer at Ahrefs

Reviewed by

Your content can be great and still never make it into an AI answer.

Every AI optimization tactic assumes AI systems can reach your pages, read them, and work out who you are. None of that is a given. Around 5.89% of websites block GPTBot outright, ChatGPT appears to abandon pages that take longer than about two seconds to respond, and most AI crawlers don’t run JavaScript at all.

The good news is that this is the most fixable part of AEO. Nothing here wins you a citation on its own, but any one of these problems can quietly cost you every citation you’d otherwise have earned.

Keep learning on YouTube


Part 1

Don’t block AI crawlers in robots.txt

None of the tips in this guide matter if AI crawlers can’t read your pages. Blocked pages can’t be cited, and it often happens by accident.

We analyzed around 140 million websites and found GPTBot is now the most blocked AI bot (5.89% of sites), with ClaudeBot close behind.

Cloudflare defaults to blocking AI platforms, so check your CDN hasn’t quietly opted you out.

Diagram: a robots.txt file blocks a crawler, drawn as a pink spider holding up an orange “a”, so the crawler can’t reach the web page and the page can’t be cited in AI search.

“Allowing AI crawlers” isn’t one decision, because each platform runs several bots that do different jobs.

OpenAI alone runs three:

  • GPTBot collects content for training
  • OAI-SearchBot builds the search index ChatGPT retrieves from
  • ChatGPT-User fetches a page live when someone asks about it mid-conversation

The trade-offs are different for each. Blocking GPTBot keeps your content out of the next model’s training data, which is the decision publishers are usually weighing. Blocking OAI-SearchBot removes you from the index ChatGPT searches, which is what costs you citations and referral traffic today. Plenty of sites block the second by accident while intending only the first.

The quickest way to see where you stand is a tool like AI Crawler Access Checker, which reads your robots.txt and tells you which crawlers are allowed through.

AI Crawler Access Checker results for ahrefs.com, loaded from ahrefs.com/robots.txt. All 27 listed crawlers, from Amazonbot to anthropic-ai, show as Allowed, and the GPTBot and OpenAI entries are ringed.

To see which bots are actually turning up, rather than which ones are permitted, check Ahrefs’ Bot Analytics. It tracks visits across 12 bot categories, including a dedicated AI filter, using server-side data through Cloudflare, so it works without adding any JavaScript to your site.


Part 2

Don’t hide content behind heavy JavaScript

AI crawlers tend to read the raw HTML and mostly don’t render JavaScript the way a browser does.

That means anything that only appears after client-side JS can be invisible to them.

Play it safe and stick to server-side rendering to give your page the best chance of AI visibility.


Part 3

Improve your page speed

Slow loading pages can negatively impact your performance in traditional search, but if anything, it seems to have an even greater penalty in answer engines. Page speed seems to help decide whether you make it into the model’s context window at all.1

Check your page speed (known as time-to-first-byte) in Site Audit:

1
Head to the Page Explorer report
2
Set up an advanced filter for: Time to first byte (ms) > 200
3
Then sort by organic traffic to find your most important content that may need to be optimized
Ahrefs’ Site Audit Page explorer with numbered badges 1 to 3: on the Page explorer heading, on an advanced filter where current Time to first byte (ms) is greater than 200, and on the Organic traffic column. The filter matches 162 results.

Part 4

Don’t bother with llms.txt

llms.txt is a proposed standard for a Markdown file at your root domain (e.g. website.com/llms.txt) that tells LLMs where your best content lives: documentation, policies, product taxonomies. In principle, it’s like robots.txt for language models.

But in practice, no major LLM provider has committed to reading it. Not OpenAI, not Anthropic, not Google.

We checked what actually happens when a website adds llms.txt. Across 137,000 domains using Ahrefs Web Analytics, 28% publish an llms.txt file. Of the ~38,000 with a valid one, 97% received zero requests for it in the month we measured.

97% of llms.txt files are never requested

Ahrefs study of 137,000 domains using Ahrefs Web Analytics.

Bar chart of domains: 137K in total, 38K with a valid llms.txt, and 1.1K whose file got any requests. A bracket over the last bar marks 97% of files getting zero requests.

Of the small share of files that did get fetched, 96% of requests came from bots. Slackbot, the link-preview bot, fetched llms.txt more often than PerplexityBot did. Put simply, the most important AI search engines do not seem to use it.

Google has been explicit about this. Its guide on optimizing for generative AI features has a section literally titled “mythbusting” that tells site owners machine-readable files like llms.txt aren’t needed to appear in generative AI search. John Mueller went further, comparing it to the old keywords meta tag:

Quotation marks

AFAIK none of the AI services have said they’re using LLMs.TXT (and you can tell when you look at your server logs that they don’t even check for it). To me, it’s comparable to the keywords meta tag – this is what a site-owner claims their site is about … (Is the site really like that? well, you can check it. At that point, why not just check the site directly?)

John Mueller, Search Advocate, Google

So should you make one? If you publish developer documentation, it costs about ten minutes and there is a plausible use case. Mueller described it as a “temporary crutch, perhaps to save some tokens” for AI coding tools parsing docs. For everyone else, it’s effort with no measured return, plus a small downside in making your content easier for competitors to scrape wholesale.


Part 5

Fix the URLs AI hallucinates about your site

AI assistants make up links to your website, and they do it often. We checked the HTTP status of 16 million URLs cited by ChatGPT, Perplexity, Copilot, Gemini, Claude, and Mistral, and found AI assistants send visitors to 404 pages 2.87x more often than Google Search does.

ChatGPT is the worst offender by a wide margin. 1.01% of its clicked URLs return a 404, against a Google baseline of 0.15%.

404 rates by referral channel

Web Analytics data - 16,000,000 URLs analyzed.

Bar chart of the estimated 404 rate by AI referrer: ChatGPT 1.01%, Perplexity 0.31%, Copilot 0.34%, Gemini 0.21%, Claude 0.58%, and Mistral 0.12%.

Perplexity (0.87%) and Gemini (0.86%) sit almost exactly on Google’s own 0.84% rate for cited URLs. That makes sense, since both retrieve from the Google index, so they inherit its broken links rather than inventing their own.

The invented URLs get inferred from your existing patterns. A common hallucinated URL for our site was ahrefs.com/keywords, because we write about keywords constantly and the model assumed a page like that must exist. It doesn’t.

This is one of the few genuinely free wins in AEO. Someone with real intent asked about your product, got sent to your site, and hit a dead end. You’ll encounter three different types:

  • Hallucinated URLs. Pages a model invented from your URL structure.
  • Outdated URLs. Pages that existed once, remembered from training data or other stale data sources.
  • Misspelled or variant URLs. Near-misses on real pages.

For that reason, we’ve added a filter to Web Analytics that will help you find hallucinated URLs in just two clicks. If you’re looking for a simple Google Analytics alternative, free for up to 1 million events each month, check it out:

Ahrefs Web Analytics Pages report with the Exit pages dropdown ringed. In the open menu, the ringed “Possible 404” option lists pages that receive visits but have a title containing 404 or not found, which can happen when AI chatbots generate hallucinated URLs.

If a hallucinated URL keeps getting requested, the model is telling you which page your visitors expect you to have.


Part 6

Become a recognized entity

Every tactic in this guide ultimately depends on one thing: whether AI systems recognize your brand as a distinct entity.

Did you know?

An entity is any distinct, real-world thing a search engine or AI can recognize and tell apart from everything else—a specific person, place, company, product, or concept (like “Dmytro Gerasymenko”, “Ahrefs”, or “Singapore”) rather than just the words used to describe it.

An AI Overview about Ahrefs with three entities highlighted in yellow and labelled beside it: Ahrefs as a company entity, Dmytro Gerasymenko as a person entity, and Singapore as a location entity.

Google’s Knowledge Graph—a database of 54 billion entities and over 1.6 trillion facts about how they relate—is no longer just the engine behind Knowledge Panels.

It’s now the core infrastructure for AI Overviews, AI Mode, and Gemini.

If your brand isn’t clearly established as an entity, with consistent signals across your site, structured data, and authoritative third-party sources, you risk being invisible in AI answers—not just the blue links.

Schema often gets pitched as a fast AI-visibility win, and the surface correlation looks convincing: AI-cited pages are nearly three times more likely to have JSON-LD.

Schema is much more common on pages cited by AI

Ahrefs study of 6M pages (2M pages per group).

Bar chart of pages with and without JSON-LD: 18.3% of non-cited pages have it, compared to 51.6% of reference-cited pages and 53.1% of inline-cited pages.

But correlation isn’t cause. We tracked 1,885 pages that added JSON-LD against 4,000 controls and found no meaningful uplift on AI Mode or ChatGPT, and a small AI Overview decline we couldn’t pin on schema.

So is schema useless? No, but its role is slow and indirect, not a quick-win tactic.

You can check whether Google already recognizes you using Google’s Knowledge Graph API or a free checker to see what entity type it assigned you and whether the attributes are right.

For instance, Carl Hendy’s Knowledge Graph Tool shows you results for people, companies, products, places, and other entities.

Google’s Knowledge Graph Search Tool on audits.com with a search for “CNN”. It found 20 results, led by Cable News Network, Inc., typed as Corporation and Organization, followed by CNN Türk and CNN Brazil.

From there, try to earn mentions from authoritative publications and keep your facts consistent everywhere: the same name, description, and category across your site, your profiles, and third-party listings.

Then add schema where it’s natural, prioritizing the types that define your entity and its relationships:

  • Organization and sameAs schema linking to your Wikipedia, Wikidata, and social profiles

  • Person schema for your site’s authors

  • Product schema for your offerings

Tip

Find your highest-authority pages and make sure those carry clean schema markup.

You can do this by heading to the Page Explorer report in Site Audit, and sorting by “Page Rating”.

Page Rating (PR) shows your URL’s internal and external backlink profile strength relative to other URLs included in your crawl.

Then check the “Schema items” column to audit your existing schema, and the “Structured data issues” column to catch missing or invalid structured data.

Ahrefs’ Site Audit Page explorer sorted by PR, with the PR column and the Schema items and Structured data issues columns ringed. The homepage has a PR of 77 and Organization schema with a Schema.org validation error, and the pricing page carries FAQPage and Organization schema with the same error.

Part 7

Prepare for agent-to-agent marketing

I’ll leave you with this final thought: a growing number of the “readers” you’re attempting to influence are actually AI agents browsing and buying on your audience’s behalf.

In fact in June 2026, Cloudflare CEO, Matthew Prince, revealed that agentic (bot/crawler/agent) traffic had surpassed the 50% mark for the first time in the history of the internet.2, 3

An X post by Cloudflare CEO Matthew Prince from June 3, 2026, reading: “Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots have now passed human traffic online for the first time in the Internet’s history.” It links to Cloudflare Radar’s Traffic Worldwide page.

Far from being an edge case or a novelty, searches for “agentic AI” have jumped 84x in three years, as the cost of running an agent has fallen from dollars to cents.

Monthly US search volume

Keywords: “agentic AI” and “what is agentic AI”

Line chart of monthly US search volume from 2023 to 2026. “agentic AI” (blue) climbs from almost nothing in 2024 to 122,175, while “what is agentic AI” (pink) peaks in mid-2025 and ends at 14,206.

Source: Ahrefs via Agent A

You don’t need to rebuild your website solely for AI agents just yet. But the old advice about creating content solely for humans and not worrying about robots probably doesn’t hold up quite as much as it used to.

This chapter is about influencing the AI the human trusts, and agent-to-agent marketing is that—but taken one step further.

Portrait of Louise Linehan

Louise is a Content Marketer at Ahrefs. Over the past ten years, she has held senior content positions at SaaS brands: Pi Datametrics, BuzzSumo, and Cision. During this time, she has published hundreds of blog posts, spearheaded industry-leading research, and developed expert-led webinar programs.

Louise first learned the ropes of SEO during her time at enterprise SEO company, Pi. Within weeks of joining, she had published a piece on cannibalization that caught the attention of the biggest name in SEO at the time: Rand Fishkin. Since then, she has written on topics ranging from manipulative link building, to content syndication, and search demand forecasting. In her spare time, she has deepened her practical understanding of SEO by building websites for two local businesses.

You can read Louise’s words in industry publications like Search Engine Watch, The Drum, The Telegraph, and Spin Sucks.

Reviewed by

Explore all AEO guides

/01

What Is Answer Engine Optimization (AEO)?

Start here. What AEO is, how it differs from SEO, and the seven steps to getting started.

/02

How Answer Engines Work

What actually happens between a prompt and an answer — and where your brand can enter the process.

/03

How to Track AI Visibility

You can’t optimize what you can’t see. Seven ways to measure where your brand shows up in AI answers.

/04

How to Choose the Best Prompts to Track

Track clusters of prompts, not one-off responses. Eleven places to find the queries worth monitoring.

/05

How to Optimize Content for Answer Engines

Ranking gets you considered. This is how you get cited — the topics, formats, and on-page changes that work.

/06

How to Monitor and Win Brand Mentions in AI Answers

AI recommends the brands it has seen discussed elsewhere. Here’s how to monitor those mentions and earn more of them.

/07

How to Optimize Your Website for Answer Engines

The technical fixes that decide whether AI can reach your pages, read them, and work out who you are.