Content Signals Policy: What robots.txt and llms.txt Mean for Your Content Strategy
AI is crawling your site right now. The question is whether you're telling it what to do — or letting it decide for itself. Here's how robots.txt and llms.txt give you back control.
John DoeCEO3 min readAI is crawling your site right now. The question is whether you're telling it what to do — or letting it decide for itself.
GPTBot (OpenAI's crawler) has grown from 5% of all AI crawler traffic to 30% in a single year. PerplexityBot saw a 157,490% increase in raw requests over the same period. These aren't edge cases. They're the new normal.
That's where Content Signals Policy comes in — and why two files, robots.txt and llms.txt, are becoming essential tools for any content team that wants to stay in control.
What Is a Content Signals Policy?
A Content Signals Policy is your brand's deliberate stance on how AI systems can access, use and cite your content. It answers three questions: Which AI crawlers are allowed to access your site? Can they use your content to train their models? How should they represent your content when answering user queries?
Without a policy, you're not neutral — you're opted in by default. Silence is permission.
robots.txt: The First Line of Defence
robots.txt has been around since 1994. It's a plain text file that sits in your website's root directory and tells crawlers what they can and can't access. Until recently, it was mostly a conversation between your site and Googlebot. That's changed. Now it's the primary tool for managing a growing list of AI bots — each with different purposes, different compliance records and different implications for your content.
The Google-Extended Exception
Google-Extended doesn't behave like a traditional crawler. It's a control token, not a user agent. When you block it, you're not blocking Googlebot — you're telling Google not to use your content to train Gemini. Your site still gets indexed normally. Your AI Overviews visibility is unaffected. This single directive keeps your content out of Google's AI training pipeline while leaving your search presence intact. It's one of the most important and underused signals available.
llms.txt: The Other Side of the Equation
While robots.txt is about restriction, llms.txt is about opportunity. Proposed by Jeremy Howard in September 2024, llms.txt is a Markdown file that lives at the root of your site. Its purpose is to help AI models understand your content — quickly, accurately and in context — at the point when a user is actively searching for information.
Generative Engine Optimisation — GEO — is how you get cited by AI tools like ChatGPT, Claude and Perplexity, not just ranked by Google. A well-structured llms.txt file makes your site dramatically easier for AI to process. It's one of the most practical GEO moves you can make right now.
The Strategic Decision: Block, Allow or Optimise?
Block AI training crawlers if you're concerned about your content being absorbed into models without attribution. Allow AI search crawlers if you want referral traffic and citations from tools like Perplexity and ChatGPT. Optimise with llms.txt if your goal is visibility in AI-generated answers.
What This Means for Your Content Programme
Right now, only around 14% of major domains have any specific AI directives in their robots.txt. That means most of your competitors have no policy at all. Getting this right now is a genuine competitive advantage. At Content Gurus, we build robots.txt and llms.txt strategy into every content programme we run. Contact Content Gurus today for a free content audit — we'll tell you exactly what the bots are seeing, and what to do about it.