
Content Marketer at Ahrefs
Your content can be great and still never make it into an AI answer.
Every AI optimization tactic assumes AI systems can reach your pages, read them, and work out who you are. None of that is a given. Around 5.89% of websites block GPTBot outright, ChatGPT appears to abandon pages that take longer than about two seconds to respond, and most AI crawlers don’t run JavaScript at all.
The good news is that this is the most fixable part of AEO. Nothing here wins you a citation on its own, but any one of these problems can quietly cost you every citation you’d otherwise have earned.
Keep learning on YouTube
None of the tips in this guide matter if AI crawlers can’t read your pages. Blocked pages can’t be cited, and it often happens by accident.
We analyzed around 140 million websites and found GPTBot is now the most blocked AI bot (5.89% of sites), with ClaudeBot close behind.
Cloudflare defaults to blocking AI platforms, so check your CDN hasn’t quietly opted you out.
“Allowing AI crawlers” isn’t one decision, because each platform runs several bots that do different jobs.
OpenAI alone runs three:
The trade-offs are different for each. Blocking GPTBot keeps your content out of the next model’s training data, which is the decision publishers are usually weighing. Blocking OAI-SearchBot removes you from the index ChatGPT searches, which is what costs you citations and referral traffic today. Plenty of sites block the second by accident while intending only the first.
The quickest way to see where you stand is a tool like AI Crawler Access Checker, which reads your robots.txt and tells you which crawlers are allowed through.

To see which bots are actually turning up, rather than which ones are permitted, check Ahrefs’ Bot Analytics. It tracks visits across 12 bot categories, including a dedicated AI filter, using server-side data through Cloudflare, so it works without adding any JavaScript to your site.
AI crawlers tend to read the raw HTML and mostly don’t render JavaScript the way a browser does.
That means anything that only appears after client-side JS can be invisible to them.
Play it safe and stick to server-side rendering to give your page the best chance of AI visibility.
Slow loading pages can negatively impact your performance in traditional search, but if anything, it seems to have an even greater penalty in answer engines. Page speed seems to help decide whether you make it into the model’s context window at all.1
Check your page speed (known as time-to-first-byte) in Site Audit:
llms.txt is a proposed standard for a Markdown file at your root domain (e.g. website.com/llms.txt) that tells LLMs where your best content lives: documentation, policies, product taxonomies. In principle, it’s like robots.txt for language models.
But in practice, no major LLM provider has committed to reading it. Not OpenAI, not Anthropic, not Google.
We checked what actually happens when a website adds llms.txt. Across 137,000 domains using Ahrefs Web Analytics, 28% publish an llms.txt file. Of the ~38,000 with a valid one, 97% received zero requests for it in the month we measured.
97% of llms.txt files are never requested
Ahrefs study of 137,000 domains using Ahrefs Web Analytics.

Of the small share of files that did get fetched, 96% of requests came from bots. Slackbot, the link-preview bot, fetched llms.txt more often than PerplexityBot did. Put simply, the most important AI search engines do not seem to use it.
Google has been explicit about this. Its guide on optimizing for generative AI features has a section literally titled “mythbusting” that tells site owners machine-readable files like llms.txt aren’t needed to appear in generative AI search. John Mueller went further, comparing it to the old keywords meta tag:
AFAIK none of the AI services have said they’re using LLMs.TXT (and you can tell when you look at your server logs that they don’t even check for it). To me, it’s comparable to the keywords meta tag – this is what a site-owner claims their site is about … (Is the site really like that? well, you can check it. At that point, why not just check the site directly?)
John Mueller, Search Advocate, Google
So should you make one? If you publish developer documentation, it costs about ten minutes and there is a plausible use case. Mueller described it as a “temporary crutch, perhaps to save some tokens” for AI coding tools parsing docs. For everyone else, it’s effort with no measured return, plus a small downside in making your content easier for competitors to scrape wholesale.
AI assistants make up links to your website, and they do it often. We checked the HTTP status of 16 million URLs cited by ChatGPT, Perplexity, Copilot, Gemini, Claude, and Mistral, and found AI assistants send visitors to 404 pages 2.87x more often than Google Search does.
ChatGPT is the worst offender by a wide margin. 1.01% of its clicked URLs return a 404, against a Google baseline of 0.15%.
404 rates by referral channel
Web Analytics data - 16,000,000 URLs analyzed.

Perplexity (0.87%) and Gemini (0.86%) sit almost exactly on Google’s own 0.84% rate for cited URLs. That makes sense, since both retrieve from the Google index, so they inherit its broken links rather than inventing their own.
The invented URLs get inferred from your existing patterns. A common hallucinated URL for our site was ahrefs.com/keywords, because we write about keywords constantly and the model assumed a page like that must exist. It doesn’t.
This is one of the few genuinely free wins in AEO. Someone with real intent asked about your product, got sent to your site, and hit a dead end. You’ll encounter three different types:
For that reason, we’ve added a filter to Web Analytics that will help you find hallucinated URLs in just two clicks. If you’re looking for a simple Google Analytics alternative, free for up to 1 million events each month, check it out:
If a hallucinated URL keeps getting requested, the model is telling you which page your visitors expect you to have.
Every tactic in this guide ultimately depends on one thing: whether AI systems recognize your brand as a distinct entity.
Did you know?
An entity is any distinct, real-world thing a search engine or AI can recognize and tell apart from everything else—a specific person, place, company, product, or concept (like “Dmytro Gerasymenko”, “Ahrefs”, or “Singapore”) rather than just the words used to describe it.

Google’s Knowledge Graph—a database of 54 billion entities and over 1.6 trillion facts about how they relate—is no longer just the engine behind Knowledge Panels.
It’s now the core infrastructure for AI Overviews, AI Mode, and Gemini.
If your brand isn’t clearly established as an entity, with consistent signals across your site, structured data, and authoritative third-party sources, you risk being invisible in AI answers—not just the blue links.
Schema often gets pitched as a fast AI-visibility win, and the surface correlation looks convincing: AI-cited pages are nearly three times more likely to have JSON-LD.
Schema is much more common on pages cited by AI
Ahrefs study of 6M pages (2M pages per group).

But correlation isn’t cause. We tracked 1,885 pages that added JSON-LD against 4,000 controls and found no meaningful uplift on AI Mode or ChatGPT, and a small AI Overview decline we couldn’t pin on schema.
So is schema useless? No, but its role is slow and indirect, not a quick-win tactic.
You can check whether Google already recognizes you using Google’s Knowledge Graph API or a free checker to see what entity type it assigned you and whether the attributes are right.
For instance, Carl Hendy’s Knowledge Graph Tool shows you results for people, companies, products, places, and other entities.

From there, try to earn mentions from authoritative publications and keep your facts consistent everywhere: the same name, description, and category across your site, your profiles, and third-party listings.
Then add schema where it’s natural, prioritizing the types that define your entity and its relationships:
Organization and sameAs schema linking to your Wikipedia, Wikidata, and social profiles
Person schema for your site’s authors
Product schema for your offerings
Tip
Find your highest-authority pages and make sure those carry clean schema markup.
You can do this by heading to the Page Explorer report in Site Audit, and sorting by “Page Rating”.
Page Rating (PR) shows your URL’s internal and external backlink profile strength relative to other URLs included in your crawl.
Then check the “Schema items” column to audit your existing schema, and the “Structured data issues” column to catch missing or invalid structured data.
I’ll leave you with this final thought: a growing number of the “readers” you’re attempting to influence are actually AI agents browsing and buying on your audience’s behalf.
In fact in June 2026, Cloudflare CEO, Matthew Prince, revealed that agentic (bot/crawler/agent) traffic had surpassed the 50% mark for the first time in the history of the internet.2, 3

Far from being an edge case or a novelty, searches for “agentic AI” have jumped 84x in three years, as the cost of running an agent has fallen from dollars to cents.
You don’t need to rebuild your website solely for AI agents just yet. But the old advice about creating content solely for humans and not worrying about robots probably doesn’t hold up quite as much as it used to.
This chapter is about influencing the AI the human trusts, and agent-to-agent marketing is that—but taken one step further.

Louise is a Content Marketer at Ahrefs. Over the past ten years, she has held senior content positions at SaaS brands: Pi Datametrics, BuzzSumo, and Cision. During this time, she has published hundreds of blog posts, spearheaded industry-leading research, and developed expert-led webinar programs.
Louise first learned the ropes of SEO during her time at enterprise SEO company, Pi. Within weeks of joining, she had published a piece on cannibalization that caught the attention of the biggest name in SEO at the time: Rand Fishkin. Since then, she has written on topics ranging from manipulative link building, to content syndication, and search demand forecasting. In her spare time, she has deepened her practical understanding of SEO by building websites for two local businesses.
You can read Louise’s words in industry publications like Search Engine Watch, The Drum, The Telegraph, and Spin Sucks.
Start here. What AEO is, how it differs from SEO, and the seven steps to getting started.
What actually happens between a prompt and an answer — and where your brand can enter the process.
You can’t optimize what you can’t see. Seven ways to measure where your brand shows up in AI answers.
Track clusters of prompts, not one-off responses. Eleven places to find the queries worth monitoring.
Ranking gets you considered. This is how you get cited — the topics, formats, and on-page changes that work.
AI recommends the brands it has seen discussed elsewhere. Here’s how to monitor those mentions and earn more of them.
The technical fixes that decide whether AI can reach your pages, read them, and work out who you are.