Free Robots.txt Generator — Create, Validate & Block AI Bots
Generate robots.txt files with 25+ AI bot controls, 12 CMS presets, and real-time validation. UtilHub's Robots.txt Generator gives you complete control over how search engine crawlers and AI bots interact with your website — all from a visual interface that runs entirely in your browser. Nothing is uploaded to our servers, and your configuration data never leaves your device.
As of March 2026, over 5.6 million websites block OpenAI's GPTBot and 5.8 million block Anthropic's ClaudeBot. AI bot traffic quadrupled during 2025, with HUMAN Security reporting a 6,900% year-over-year increase in verified AI agent traffic. Our generator includes all current AI crawlers — training bots, search bots, and user-triggered bots — organized by purpose so you can make informed blocking decisions.
How to use Robots.txt Generator
- Choose your setup method — Select "Create from Scratch" to build custom rules, "CMS Templates" for WordPress, Shopify, Joomla, and 9 other ready-made configurations, or "Validate & Test" to check an existing robots.txt file. CMS templates include pre-configured rules tested with 12 popular platforms.
- Configure bot access rules — Set allow or disallow rules for search engine bots (Googlebot, Bingbot, Yandex) and choose from 25+ AI crawlers organized by purpose: training bots (GPTBot, ClaudeBot), search bots (OAI-SearchBot, PerplexityBot), and user-triggered bots. Use quick-action buttons like "Block All AI Training" or "SEO Recommended" for instant configuration. Add restricted directories like /admin/, /search/, or /cart/.
- Add sitemap and review output — Enter your XML sitemap URL so search engines can discover all your pages. Add a crawl-delay only if Bingbot is hitting your server too hard — Google has never supported the directive and Yandex stopped honouring it in 2018. Review the syntax-highlighted output in real-time as you make changes — every edit updates the generated file instantly.
- Validate, download, and upload — Use the built-in validator to check for errors before deploying. Copy the generated robots.txt to your clipboard or download it as a file. Upload it to your website's root directory — it must be accessible at https://yoursite.com/robots.txt. Then check it in Search Console under Settings → robots.txt, the report that replaced the retired robots.txt Tester, and use URL Inspection to confirm whether a specific URL is blocked.
Features
- 25+ AI Bot Controls — Block or allow AI training crawlers (GPTBot, ClaudeBot, Google-Extended), AI search crawlers (OAI-SearchBot, PerplexityBot), and user-triggered bots, organized by purpose with clear explanations.
- 12 CMS Presets — Ready-made templates for WordPress, WooCommerce, Shopify, Joomla, Drupal, Magento, Next.js, Laravel, Angular, Wix, Squarespace, and Webflow with platform-specific best practices.
- Real-Time Validation — Built-in syntax checker catches errors, warnings, and RFC 9309 compliance issues before you deploy. Color-coded results make problems obvious.
- URL Path Tester — Test any URL path against your robots.txt rules to verify whether it is allowed or blocked, with the matching rule highlighted.
- Syntax-Highlighted Preview — Live preview with color-coded directives (User-agent in green, Disallow in red, Allow in blue) updates instantly as you configure rules.
- Quick Action Presets — One-click configurations including "Block All AI Bots", "SEO Recommended", "Allow Only Search Engines", and "Block Everything" for instant setup.
Frequently Asked Questions
What is a robots.txt file and why do I need one?
A robots.txt file is a plain text file placed in your website's root directory that tells search engine crawlers and bots which pages they can and cannot access. It follows the Robots Exclusion Protocol, formalized as RFC 9309 in September 2022. While search engines will crawl your site without one, a robots.txt file helps manage crawl budget — the limited number of pages Google crawls per visit. By blocking non-essential pages like admin panels, internal search results, and duplicate content, you ensure crawlers focus on your important pages, improving indexing speed and SEO performance.
How do I block AI crawlers like ChatGPT and Claude from scraping my content?
Add specific User-agent rules for each AI bot you want to block. As of March 2026, the major AI crawlers are: GPTBot (OpenAI/ChatGPT training), OAI-SearchBot (ChatGPT Search indexing), ClaudeBot (Anthropic/Claude training), Google-Extended (Gemini AI training), CCBot (Common Crawl, used for AI datasets), PerplexityBot (Perplexity AI), Meta-ExternalAgent (Meta/Llama), Bytespider (ByteDance/TikTok AI), and DeepSeekBot (DeepSeek). For each training bot, add "User-agent: [bot-name]" followed by "Disallow: /" to block your entire site. Over 5.6 million websites now block GPTBot. Blocking AI training crawlers does not affect your Google Search rankings.
What is the difference between Disallow and Noindex?
Disallow in robots.txt prevents crawlers from accessing a page, but it does not guarantee the page won't appear in search results — Google can still index a URL based on external links without crawling it. The noindex meta tag tells search engines not to show a page in results, but crawlers must first access the page to read the directive. For complete removal from search results, do not block the page with robots.txt (so crawlers can read the noindex tag) and add the noindex directive to the page itself. Google officially deprecated the unofficial robots.txt noindex directive on September 1, 2019.
What is crawl-delay and should I use it?
Crawl-delay tells bots to wait a specified number of seconds between requests to your server. It is not part of RFC 9309 and support has been shrinking. Google has never honoured it, and the Search Console crawl rate limiter that used to be the alternative was retired on 8 January 2024 — Googlebot now sets its own rate from your server's response times and error rates, so the way to slow it down is to fix slow responses or return 503 or 429 for a short period, and to use Google's Googlebot report form if crawling is still excessive. Bing still respects values of 1-30 seconds. Yandex stopped taking the directive into account on 22 February 2018 and uses the crawl rate setting in Yandex Webmaster instead. For most websites, crawl-delay is unnecessary and can actually slow down indexing. Only use it if your server has limited resources and bot traffic causes performance issues — typically shared hosting with high-traffic sites.
Where do I upload the robots.txt file?
The robots.txt file must be placed in your website's root directory, accessible at https://yoursite.com/robots.txt. For traditional hosting, upload via FTP/SFTP to the public_html folder. In WordPress, use Yoast SEO (Tools → File Editor) or Rank Math (General Settings → Edit robots.txt). For Shopify, edit the robots.txt.liquid template in your theme code. On Cloudflare Pages, place it in your public/ source directory. For Angular 22 with SSR, add it to the assets array in angular.json. After uploading, verify by visiting yoursite.com/robots.txt in your browser, then test with Google Search Console.
Can robots.txt protect private or sensitive content?
No. Robots.txt is not a security mechanism — it is publicly readable by anyone at yoursite.com/robots.txt, which actually reveals which directories you consider sensitive. Malicious bots ignore robots.txt entirely, and as of 2026, about 13% of AI bot requests also ignore it. For private content, use proper authentication, server-side access controls (.htaccess, firewall rules), password protection, or IP whitelisting. Robots.txt should only be used for crawl management — controlling how legitimate bots interact with your public content.
Does blocking AI bots affect my Google Search rankings?
No. Blocking AI training crawlers like GPTBot, ClaudeBot, CCBot, or Google-Extended has no effect on your Google Search rankings. Googlebot (the crawler responsible for search indexing) is completely separate from Google-Extended (which controls AI training data use). However, blocking AI search crawlers like OAI-SearchBot or PerplexityBot means your content will not appear in those AI search products, which may reduce your overall traffic from AI-powered search.
What is RFC 9309 and why does it matter for robots.txt?
RFC 9309 is the formal Internet Engineering Task Force (IETF) standard for robots.txt, published in September 2022. Before RFC 9309, robots.txt was a de facto convention with no official specification, leading to inconsistent implementations. The standard formalizes User-agent, Disallow, and Allow as the only recognized directives, standardizes wildcard patterns (* and $), codifies error handling (4xx means full access, 5xx means full disallow), sets a minimum file size of 500 KiB, and specifies that the most specific (longest path) rule wins when conflicts occur. Crawl-delay, Sitemap, Noindex, Host, and Clean-param are explicitly excluded from the standard, though Sitemap remains widely supported as an extension.