Key takeaways
- GPTBot, PerplexityBot and Google-Extended are the only crawlers that feed AI answer engines directly, so blocking any of them removes your business from that engine's citations.
- A single Disallow line added by a security plugin or developer default can erase a business from ChatGPT and Perplexity answers with no error message and no way to know it happened.
- Robots.txt is a voluntary instruction crawlers can choose to ignore, not an access restriction, so it never protects pricing pages or proprietary content from a competitor's browser.
- AI models only cite businesses they can verify, so consistent name-address-phone data, service-specific pages and complete schema markup matter as much as crawler access itself.
- Checking yoursite.com/robots.txt and testing what a real prospect would ask ChatGPT or Perplexity monthly is the only reliable way to know whether your business is currently citable.
Should you block AI crawlers like GPTBot on your website? For most service businesses, no — and doing it reflexively is one of the fastest ways to disappear from ChatGPT, Perplexity and Google AI Overviews before those channels ever send you a lead. The decision isn't about privacy or control; it's about which bots feed the answer engines your customers already use, and which ones just scrape your content for someone else's product.
For a contractor running a $3,000-a-month ad budget with a two-person sales team, the answer channels outside paid search are not optional extras — they're free distribution you can't afford to close off by accident. A single disallow line in robots.txt, added by a plugin default or a well-meaning developer, can remove you from every AI-generated answer in your category with no error message and no way to know it happened until a competitor shows up in your place instead.
Blocking GPTBot removes you from ChatGPT's answers
Blocking GPTBot in robots.txt tells OpenAI's crawler not to read your site, and that has one direct consequence: your business stops showing up when someone asks ChatGPT for a recommendation in your category. If a homeowner asks "who does roofing repair near me" or "best HVAC company for a same-day quote," ChatGPT can only cite businesses it has crawled and can verify. Block the crawler, and you've opted out of that answer pool entirely — not just for today's query, but for every future one, since these models are retrained and refreshed on crawl data over time. For a contractor or local service business competing for a handful of high-intent searches a month, that's not a neutral choice. It's a lead source turned off at the source.
The technical trigger is a two-line block in robots.txt — a User-agent: GPTBot line followed by Disallow: / — and it's often added without a business decision behind it, buried in a security plugin's default configuration or a developer's blanket "block all bots" habit carried over from an unrelated project. For a two-person sales team that can't outbid national competitors on Google Ads, an unpaid citation inside a ChatGPT answer is exactly the kind of lead source that never shows up in a media budget line item but still closes real jobs.
The three crawlers that actually determine AI visibility
GPTBot (OpenAI/ChatGPT), PerplexityBot (Perplexity), and Google-Extended (Google AI Overviews and Gemini) are the crawlers that feed answer engines directly — block any of them and you lose eligibility for citations from that specific model. These aren't the only AI-related bots hitting your server logs, but they're the ones tied to a real, revenue-relevant answer surface. Contrast that with bots like Bytespider (TikTok's parent company) or CCBot (Common Crawl, used to train many open models indirectly) — blocking those has little effect on whether you get cited in a customer-facing answer today, since they don't power a consumer-facing assistant your buyers are querying directly.
The practical distinction: ask whether a real prospect could type a question into that specific assistant and get your business named in the answer. If yes, that crawler needs access. If the bot only feeds a training pipeline with no direct query interface your customers use, blocking it costs you little and protects you from unrestricted scraping. Each crawler respects its own user-agent line, so you can permit one and block another in the same file — a User-agent: GPTBot block with Allow: / can sit directly above a User-agent: Bytespider block with Disallow: / without conflict. Checking your server logs for these exact user-agent strings, rather than guessing from a plugin's dashboard label, is the only reliable way to confirm which crawlers are actually reaching your pages versus which ones your hosting platform is silently blocking upstream.
Entity clarity matters more than crawler access alone
Allowing GPTBot to crawl your site is necessary but not sufficient — the bigger lever is whether your site gives these models a clear, structured answer once they arrive. AI models cite businesses they can verify: consistent name, address and phone across your site and directory listings, service pages that state exactly what you do and where, and schema markup that removes ambiguity about who you are. A site that's crawlable but vague — generic "About Us" copy, no clear service area, no structured data — gets crawled and still doesn't get cited, because the model can't extract a confident answer from it.
In practice, that means LocalBusiness schema markup with your name, address, phone number and service categories filled in completely — not left as placeholder fields inherited from a theme template — plus service pages written for each specific job you do, rather than one generic "Services" page listing five offerings in a bulleted list. A roofing contractor with separate pages for storm damage repair, roof replacement and gutter installation gives an AI model three distinct, citable answers instead of one vague one it has to guess at.
This is the same foundation local SEO and Google Business Profile work has always required, extended to a new set of consumers: language models instead of just the Google index. If your GBP listing, website copy and directory citations already disagree on your hours or service area, that inconsistency confuses AI crawlers exactly the way it confuses human searchers — except now it costs you citations in three answer engines instead of one search results page.
When blocking a crawler is the correct call
Blocking makes sense in a narrow set of cases: crawlers you can't identify hitting your server at high volume, bots with no attached consumer product, or traffic that's degrading site performance without any citation upside. If your server logs show repeated requests from an unfamiliar user-agent string consuming bandwidth with no plausible connection to a customer-facing assistant, blocking it in robots.txt is reasonable server hygiene — the same logic as blocking any scraper that offers no return traffic or citation value. The mistake is applying that same instinct to GPTBot, PerplexityBot or Google-Extended, which do carry direct citation upside.
In practice this looks like an unfamiliar user-agent making thousands of requests a day against a site that gets a few hundred human visitors, with no corresponding referral traffic or brand mention anywhere you can find it. That pattern is a server-load problem, not an AI-visibility problem, and it's worth fixing regardless of what the bot does with the data — but it's a narrow case, not a default posture to apply to every crawler you don't recognize.
Quick test before you block anything
Robots.txt is not a privacy control
A common misconception is that blocking AI crawlers protects proprietary information or prevents competitors from seeing your pricing and process pages — it doesn't. Robots.txt is a voluntary instruction that compliant crawlers honor, not an access restriction; anything already indexed by Google or visible to a human visitor is not protected by it. If you have genuinely sensitive content (client-specific pricing, internal tools, draft pages), the fix is authentication or removing it from public URLs, not a robots.txt directive that a competitor's browser ignores completely while also blocking the AI crawlers that could have sent you a customer.
Disallow and noindex are also frequently confused: a robots.txt disallow tells compliant crawlers not to request a page, while a noindex meta tag tells search engines not to list a page they've already read — neither one restricts who can view a URL if they already have the link. For a business worried about a competitor seeing internal pricing sheets or proposal templates, the only real fix is moving that content behind a login or off public URLs entirely, not adjusting a file that well-behaved crawlers simply skip past.
How to check whether AI models already recommend you
Ask ChatGPT, Perplexity and Google AI Overviews the exact questions a prospect would type — "best [your service] in [your city]," "who should I hire for [specific job]" — and note whether your business appears, what's said about you, and which competitors show up instead. Run this monthly rather than once, since answer engines update their citation patterns as they recrawl, and a business that's absent in January can show up by March purely from crawler access and content changes, with no other input.

Keep a simple log — a spreadsheet with the date, the exact question asked, and whether your business appeared — so you can see the trend rather than relying on memory of one good or bad result. A part-time marketer running this check for fifteen minutes a month generates more useful signal than a one-time audit, because it's the trend line, not a single snapshot, that tells you whether your AEO work is actually moving the needle.
If you're not appearing, the fix usually isn't a single blocked bot — it's a combination of thin service pages, missing schema, and inconsistent NAP data across the web. Treating AEO as a checklist item you configure once in robots.txt misses that it depends on the same content and structure work behind traditional rankings, applied to a different set of readers. You can see how this plays out across a full site in our work, or get in touch if you want a specific answer on whether your site is currently crawlable and citable by the models your customers are already asking.
Verify your robots.txt isn't accidentally blocking the wrong bots
Check your current robots.txt file directly at yoursite.com/robots.txt and confirm GPTBot, PerplexityBot and Google-Extended aren't disallowed — many site builders and security plugins add blanket AI-bot blocks by default without telling you. This is worth checking even if you never intentionally configured anything: some hosting platforms and security plugins ship with "block AI scrapers" turned on as a default privacy feature, which silently removes you from citation eligibility before you've made any deliberate decision about it.
It's also worth checking one layer deeper than the robots.txt file itself: some content delivery networks and security services block AI crawlers at the network level — returning an error page instead of your content — even when robots.txt itself allows them. A quick way to catch this is to load your homepage as GPTBot would, using a user-agent switcher or asking someone technical to run a single command-line request, and confirm the page returns your actual content rather than a blocked-access page. A five-minute check across both layers can be the difference between being crawlable and invisible to the exact tools your next customer is using to find you.
Prefer it done for you? This playbook is our SEO engine: see how we run it for clients →
Frequently asked questions.
Should I block GPTBot on my website?
For most service businesses, no. GPTBot feeds ChatGPT's answers directly, so blocking it removes your business from every AI-generated recommendation in your category, not just today's search.
What's the difference between GPTBot, PerplexityBot and Google-Extended versus other AI bots?
GPTBot, PerplexityBot and Google-Extended power consumer-facing assistants that prospects query directly, so blocking them costs you citations. Bots like Bytespider or CCBot feed training pipelines with no direct query interface, so blocking those has little effect on whether you show up in an answer today.
Does blocking AI crawlers protect my pricing or proprietary pages?
No. Robots.txt is a voluntary instruction that compliant crawlers honor, not an access restriction, so anything already indexed or visible to a human visitor stays reachable. Genuinely sensitive content needs authentication or removal from public URLs instead.
How do I check if my robots.txt is blocking AI crawlers by accident?
Visit yoursite.com/robots.txt directly and confirm GPTBot, PerplexityBot and Google-Extended aren't disallowed, since many site builders and security plugins add blanket AI-bot blocks by default. It's also worth checking whether a CDN or security service blocks these crawlers at the network level even when robots.txt itself allows them.
