AI Search Crawler Access: A Practical Audit for ChatGPT, Perplexity, and Google
AI Brand Report ·
Separate search access from training controls, inspect robots and page responses, and diagnose the technical barriers that keep useful content from being retrieved.
An AI search access audit checks whether the pages you want discovered are reachable, readable, and eligible for the relevant search experience. Review crawler policy, server responses, rendered content, and indexing separately. A successful check establishes technical access, not a promise that an assistant will recommend your brand.
Marketing teams often jump from “we are missing from an answer” to “we need more content.” Sometimes the better first question is whether the existing content can be fetched at all.
Separate the policy decision from the implementation
Decide which uses of public content your organization permits before editing access rules. Search retrieval, model training, and a user-requested page fetch are not necessarily the same activity.
OpenAI documents OAI-SearchBot for search and GPTBot for content that may be used in training, with independent controls. Its ChatGPT-User agent supports user-initiated actions and is not the search-index control. See OpenAI's crawler documentation.
Perplexity identifies PerplexityBot as a search crawler and Perplexity-User as a user-requested fetcher. Its guidance also describes checking official IP ranges when configuring firewall access. See Perplexity's crawler documentation.
These distinctions help you implement an approved policy precisely. They do not justify opening private pages or disabling security controls across the site.
Audit four layers, in order
| Layer | Question | Evidence to retain |
|---|---|---|
| Policy | Is the relevant crawler allowed for this path? | Applicable robots.txt group and path rule |
| Delivery | Does the public URL return the intended page? | Status, redirects, response headers, challenge behavior |
| Content | Can a fetcher access the important answer? | Initial HTML and, where relevant, rendered HTML |
| Search eligibility | Is the page eligible in the relevant index? | Indexing inspection and exclusion reasons where available |
Do not collapse these into a single “AI ready” score. Each failure has a different owner and a different repair.
1. Inspect policy for the exact host and path
Check the production hostname and the article, product, or service URL you care about. A staging robots file or a rule for a different subdomain is not evidence about the live page.
Read the complete applicable rule group. A broad wildcard block and a more specific bot group can create behavior that is easy to misread from a single line. Have the technical owner validate the effective policy with the site's normal tooling.
Remember that robots.txt is not an authentication system. Private information belongs behind proper access controls regardless of a crawler's stated behavior.
2. Inspect delivery before changing the firewall
A browser may show the page after completing a challenge while an automated fetch receives an interstitial. Record what the fetch actually gets: an ordinary page, a redirect, an error, or a challenge.
If logs show blocks, verify the requester using the provider's documented method. A user-agent string alone is not proof of identity. Prefer narrow changes reviewed by the security owner over broad exceptions for every request claiming to be a bot.
Test the approved change on a small set of public URLs. Confirm that ordinary protections and private routes still behave as intended. Keep a rollback record containing the previous rule and the condition that would trigger restoration.
3. Check whether the answer exists in accessible content
The visible page may depend on client-side requests, tabs, or personalization. Inspect the page title, primary heading, main explanation, relevant specifications, and internal links in the content the fetcher receives.
You do not need to turn every interactive product into a static page. Provide accessible public explanations for the facts a buyer needs before logging in. For example, integration availability should not be discoverable only through an authenticated setup modal.
Verify that the canonical URL is intentional and that snippets or indexing directives have not been inherited from a preview environment. Fix the underlying page configuration rather than hiding contradictory signals in additional markup.
4. Confirm the search-specific requirements
For Google's AI Overviews and AI Mode, a page must be indexed and eligible to appear with a snippet. Google does not require special AI schema or new AI text files. See Google's AI features guidance.
That statement applies to Google's documented features. Do not extrapolate it into a universal claim about every AI assistant's retrieval system.
Use a small diagnostic sample
Choose five public pages with different roles: homepage, product page, comparison page, help article, and editorial article. Add any page that repeatedly appears in a citation audit.
For each, record an expected outcome before testing. A public comparison page should return the comparison, not an empty application shell. A discontinued product may intentionally redirect to a replacement. The audit should respect those differences.
Resolve one layer at a time and recheck. If a page becomes reachable but is still absent from answers, that does not mean the access fix failed. It means the technical barrier was removed and the remaining question concerns selection, relevance, or evidence.
Frequently asked questions
Does allowing an AI crawler guarantee a citation?
No. Access removes a potential retrieval barrier. It does not guarantee indexing, selection, citation, recommendation, or traffic.
Can I allow ChatGPT search while blocking GPTBot?
OpenAI documents OAI-SearchBot and GPTBot as independent controls for search and potential training use. Review the current official documentation before changing your policy.
Do I need a special AI schema to appear in Google AI features?
Google says no special schema or additional AI text file is required for its AI features. Normal Search eligibility and useful accessible content remain the starting point.
Connect access to actual visibility
Once access is verified, investigate what engines say about the business. Start with an AI Brand Report, then compare technical findings with content and brand evidence. Technical access and persuasive evidence solve different parts of the discovery problem.