Site owners adding AI-powered search face the same fear: will the AI answer questions using content that only some visitors should see?
The concern is legitimate, but the protection does not come from where most people look first. robots.txt, noindex and llms.txt all manage discovery by external crawlers. Your own site search is a different system: it retrieves from an index you control, at query time, for a specific user.
This guide explains how grounded answers work, where the real access boundary sits, and what you can test yourself.
Grounded AI search is safe for restricted content when retrieval is permission-scoped — robots.txt has nothing to do with it.
A grounded answer is built only from documents retrieved at query time. If the retrieval index is filtered by the requesting user's permissions on the server, restricted content never reaches the model. robots.txt and noindex govern external crawlers; they are not access control and never were.
What “grounded” actually means
A grounded answer is generated from specific documents the system fetched for that exact question — a pattern usually called retrieval-augmented generation (RAG). The model does not answer from general memory; it answers from retrieved passages, and a well-built feature shows which pages it used.
This has two practical consequences:
- The answer can only contain facts from documents the retrieval step returned.
- Whoever controls what retrieval returns controls what the AI can possibly say.
So the security question “can the AI expose restricted content?” reduces to a much older, well-understood question: does your retrieval layer enforce permissions?
Three separate access layers
Confusion comes from mixing up three independent layers. Each has its own mechanism, and none substitutes for another:
| Layer | Mechanism and what it actually does |
|---|---|
| External discovery | robots.txt, noindex, llms.txt — asks compliant crawlers what to fetch or index. Voluntary. |
| Content confidentiality | Authentication, paywalls, signed URLs — enforces who may receive the content. The only real boundary. |
| AI site search answers | Permission-scoped retrieval index — decides which documents the AI may use for each user. |
A page can be blocked in robots.txt and still fully public. A page can be allowed in robots.txt and still behind a login. The layers do not interact — which is exactly why each must be checked separately.
robots.txt is not access control
robots.txt is a public file of requests to crawlers. Compliant bots honor it; nothing technically prevents anyone else from fetching the same URL. Google’s own documentation states that robots.txt is not a mechanism for keeping a page out of reach — for that you need authentication or noindex combined with real blocking. Google: robots.txt introduction
User-agent: *
Disallow: /account/
Disallow: /internal/
# This only asks crawlers to stay out.
# /account/ must ALSO require login —
# robots.txt alone does not protect it.The same applies to noindex: it asks search engines not to index a page they can still fetch. It is a discoverability signal, not a lock.
<meta name="robots" content="noindex" />“If a URL returns content to an anonymous request, that content is public — regardless of what robots.txt, noindex or llms.txt say about it.”
Scoping what the AI can retrieve
For your own AI site search, the boundary lives in the retrieval index. The standard pattern:
- Index each document with its access level (public, customer, internal, per-account).
- At query time, filter the index by the requesting user's permissions before anything is sent to the model.
- Enforce this on the server — never in the browser, where a user could modify the filter.
- Log which documents grounded each answer, so leaks are auditable.
With this in place, an anonymous visitor’s AI search physically cannot quote a members-only page, because that page is never in the set of documents the model sees. This is the same permission model your website already uses for serving pages — extended one step further, into the index.
What about external AI assistants?
ChatGPT, Perplexity and similar systems can only ground answers in content they can fetch. If restricted content requires login, their crawlers receive a login page, not the content. Your protection there is the same authentication layer — not a crawler rule. Our guide to checking AI bot access shows how to verify what external bots actually receive, and the 12-point AI visibility checklist covers the public side.
Where structured data helps
JSON-LD structured data clarifies what a public page is about: the organization, the product, the price, the FAQ. It helps retrieval systems interpret and cite the page correctly.
One caution: structured data is part of the public HTML. Anything you put into JSON-LD on a public URL is public — including values you may consider restricted. Mark up what you want understood; never embed paywalled prices, member data or internal identifiers in markup on a public page.
What you can measure yourself
These checks need no special tools and give evidence, not impressions:
- Fetch a restricted URL without logging in (curl or an incognito window). It must return a login page or a 401/403 — not the content.
- Place a unique canary string in a restricted document, then ask your AI site search about it while logged out. The answer must not reflect it.
- Fetch your homepage with an AI bot User-Agent and compare it with the anonymous response — they should match.
- Check that no restricted values appear in JSON-LD on public pages (view source, search for the value).
- Confirm robots.txt rules match your intent for external AI crawlers — discovery only, not security.
ClearSiteScore’s free homepage scan covers the public side of this picture: which AI crawlers your robots.txt allows, what response code they receive, and whether key content is in the served HTML. It does not test your authenticated areas — those require the canary test above, run by you.
Common mistakes
| Mistake | Why it fails |
|---|---|
| “Protected” by robots.txt alone | Voluntary guideline; the URL still serves content to anyone |
| noindex on a sensitive page | The page is still fetchable; it just is not indexed |
| Hiding restricted content with JavaScript | It is in the HTML response; view source reveals it |
| Client-side filtering of the AI index | Users can modify browser-side filters; enforcement must be server-side |
| Restricted values inside public JSON-LD | Structured data is part of the public response |
FREE AI VISIBILITY CHECK
See what AI systems can actually reach on your site
Run a free homepage scan: AI crawler rules, real bot responses, served HTML, structured data and security headers. No registration, no card. Every finding shows the measured evidence behind it.
Run a free AI website auditFrequently asked questions
By separating retrieval from generation. The search index should only contain content the requesting user is allowed to see, enforced server-side at query time — not by robots.txt, which is a voluntary crawling guideline. Grounding means the answer is built only from retrieved, permitted documents, with sources shown.
References
Go deeper
Why an AI Chatbot Cannot Audit Your Website
AI VisibilityAI Search Optimization Tools in 2026: What They Check (And How to Audit Your Site for Free)
AI VisibilityHow to Check Which AI Bots Can Access Your Website
References checked on 30 September 2026. This guide describes access-control and retrieval patterns; it does not certify the behavior of any specific AI provider or search product.