What Should AI Be Allowed to Read on Your Website?

What Should AI Be Allowed to Read on Your Website?
Search your own company in an AI assistant and read what comes back. You will recognize most of the language.
Working out which page each piece came from is harder, and some of it will trace to copy you rewrote two years ago.
That summary is now part of how buyers meet you, and it was assembled by something that read your site without ever appearing in your analytics.
The advice going around is to make your website easier for AI to read.
AI should be able to read anything that explains what you sell and nothing that belongs to a named customer, and the work is sorting your site into those two piles and deciding what an agent is allowed to do once it gets there.
Why are AI crawlers reading your website more than people are?
AI crawlers read far more pages than they send visitors back to, and none of that traffic shows up in the reports your marketing team actually reviews.
Measured across Cloudflare's network in late June 2025, Anthropic's crawler made 70,900 requests for every visitor it referred back.
Other platforms look nothing like that, and Mistral sent ten referrals for every page it took over the same week, so the number worth caring about is the one attached to a specific crawler.
Nobody in your building owns that number. Your marketing lead watches sessions and conversions. Your developer watches uptime and errors.
A crawler that reads every page and refers nothing is invisible to both of them, which is how a company ends up with an AI access policy it never set. On July 1, 2025 Cloudflare began blocking AI crawlers by default for new domains, by which point more than a million of its customers had already turned blocking on.
Plenty of businesses have an answer to this sitting in a dashboard they have not opened.
Which pages do AI tools use to summarize your company?
AI summaries get assembled from whatever is public and indexable, which includes the pages your team stopped maintaining.
A service page from 2022 describing an offering you retired is still a live URL returning a 200, and pulling it out of the navigation only hid it from people.
Someone wrote that page, that person has since moved on, and deleting a page no customer has mentioned in three years is a small risk with no upside for whoever raises it.
So it stays.
Nothing on the page announces that it is out of date, because there is no field for that.
The copy reads as confidently as your current work, and a model weighing it against your newer pages has no way to tell which one you would defend today.
Then a prospect shows up to a call having read a summary built partly from work you no longer do, and you spend the first ten minutes of a sales conversation correcting a version of your company you retired on purpose.
Nothing in your analytics told you it was coming.
Go find the pages describing the business you used to have. Update them, redirect them, or take them down, and put a name next to that decision so it happens again next year.
Does llms.txt actually do anything yet?
No AI company has confirmed on record that its crawler reads llms.txt, so treat the file as cheap housekeeping and nothing more. Anthropic, Cloudflare, and Stripe all publish one, which says they expect agents to move around their own sites. Whether anyone reads yours has no published answer.
Google's John Mueller took it apart on Search Off the Record in June 2026, pointing out that a file where every publisher describes their own site cannot work as a differentiator, since you are "telling these systems, like, I have the best website ever."
He allowed one real use: a system already working inside your site can read the file to find its way around.
We get asked about llms.txt constantly, usually by someone who has been told it is the thing standing between them and AI visibility.
Writing one takes an afternoon.
That is an afternoon your stale service pages are not getting, and those are definitely being read.
Write the file if you want it, keep it accurate, and put the real effort where the summaries are actually coming from.
Which parts of your website should be public, and which belong behind a login?
The line worth drawing is between content that describes what you sell and content that belongs to one named customer.
The first kind should be easy for a machine to read correctly. A login on the second kind was already the right call before AI existed, and it is being tested more often now.
The middle is where the arguments happen. Pricing detail, technical documentation, downloadable guides, partner materials.
A gated PDF that asks for an email is a lead form with a lock icon, and an agent will supply an address to get past it. Decide whether you actually mind, because the answer is sometimes no.
For Liberty Lane we designed and built the website and admin portal for a roadside assistance program, with the public site focused on helping customers understand the program and sign up.
The platform now supports over 3,000 customers, reached in under a year. The website explains the program to anyone who asks, a crawler included, because explaining the program is what it is for.
The portal knows who signed up and what they are owed, and it asks for a login for the same reason it always did.
Four questions settle almost any page:
- Is this meant to bring in buyers, or to serve people who already bought?
- Would you email it to a stranger who asked for it?
- Does it name a customer?
- Would you be comfortable seeing it quoted in an AI answer with your company attached?
We believe that business is built on transparency and trust, and that good software is built the same way, which means a company should be able to say exactly which of its pages a stranger is allowed to read.
What happens when AI agents can submit forms on your website?
Your forms were built assuming a person was on the other side.
An agent that can fill one in and book time on your calendar changes that assumption without asking, and the first sign of it will be a lead your sales team spends a week trying to reach.
On May 6, 2025 Coinbase published x402, a standard that revives the HTTP 402 Payment Required response so a client can request a resource, get back a price, pay in stablecoin, and retry the request with the payment attached.
The payment rail is the least interesting part of it. What the standard shows is the web growing a way to answer a machine with terms attached, which is what you need when the visitor is software working for someone else.
The near-term questions are smaller and land sooner:
- Can an agent submit your contact form, and would you know?
- Can it book a consultation on your calendar?
- Can it pull a gated guide without the email gate meaning anything?
- Can it hit your API, and at what rate?
- Can it start a support request on a customer's behalf?
- Which of those should require a person before anything happens?
Rate limits, real authentication on anything touching a customer record, and a named list of actions that require a person are ordinary development work.
They are cheaper to do now than during the week your sales team is chasing phantom leads.
What should you audit on your website for AI access?
An AI access audit takes a marketing lead and a developer about a week together.
Seven things to check:
- Ask three AI assistants about your company and compare the answers to your actual positioning. Write down every claim you would correct on a call.
- Name your authoritative pages. Decide which ones are the best current explanation of what you do and who you serve.
- Find the stale ones. Pull your full list of live URLs and flag every page describing work you no longer do.
- Check your robots.txt and your CDN settings. Find out which crawlers you are already blocking, including the ones you did not choose to block.
- Review metadata and schema on the pages from step two, so the page-level signals match the business you have now.
- Test the login wall. Try reaching a customer-specific URL while logged out, on every one of them.
- Give it an owner. These rules go stale as fast as the content does, and they need a person's name on them.
How should a company decide what AI can see?
Decide AI access the way you would handle any access question, by asking what each piece of content is for and who it belongs to.
Content meant to attract buyers should be clear, current, and readable by a machine, customer data sits behind authentication, and everything in between needs a decision somebody actually made.
Open your list of live URLs and mark each one open, limited, or protected. The pages you cannot classify are the ones a buyer is most likely to hear about before you do.
Related Articles
Here are a couple related articles to view, or return back to the main page.


Check out the BIZ/DEV podcast
Our weekly tech podcast focusing on AI, our industry, the founder's journey, and more.
