Skip to content

Runs entirely in your browser. Nothing you paste leaves this page.

Free / No sign-up

Robots.txt generator and tester for AI crawlers.

Build a robots.txt with path rules, sitemaps and one-click controls for current AI crawlers, then test whether any URL is allowed for any bot, using the same matching rules Google applies.

Start from
Rule group 1
AI crawlers: tick to block the whole site

Test a URL

Allowed for Googlebot at /blog/core-web-vitals. No rule matches this path, so it is allowed.

Blocked for 8 of 18 listed crawlers: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Meta-ExternalAgent, Amazonbot, Bytespider

Matching follows RFC 9309 as Google applies it: the most specific user-agent group wins, then the longest matching path, and Allow wins a tie. robots.txt controls crawling, not indexing, and only well-behaved bots obey it.

How to use it.

  1. 01

    Start from a preset, then edit the user-agent groups and Allow or Disallow paths, and tick the AI crawlers you want to keep out.

  2. 02

    Add your sitemap URLs. The robots.txt updates live; you can also edit or paste a file directly in the text pane.

  3. 03

    Type a URL and pick a crawler to see whether it is allowed and which rule decided it, then download the file and upload it to your site root.

What it does.

Everything this tool handles, all of it inside your browser tab.

  • Rule builder with any number of user-agent groups and Allow or Disallow paths
  • One-click blocking for 14 AI crawlers and tokens, labelled by purpose: training, AI search or user fetch
  • Presets: allow all, block AI training, block all AI, block everything
  • Sitemap lines for one or more sitemaps
  • Editable output: paste an existing robots.txt to test it
  • URL tester with RFC 9309 group selection, longest-match and Allow-wins-ties logic
  • Verdict for every listed crawler at once, plus the rule that decided it
  • Flags lines crawlers will not understand
  • Download robots.txt or copy it

Worked examples.

  • Block AI training, keep AI search

    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: Google-Extended
    User-agent: Applebot-Extended
    User-agent: CCBot
    User-agent: Meta-ExternalAgent
    Disallow: /
    
    User-agent: *
    Disallow: /admin/

    OAI-SearchBot, Claude-SearchBot and PerplexityBot are not listed, so they follow the * group and can still index and cite your pages.

  • Longest match wins

    User-agent: *
    Disallow: /admin/
    Allow: /admin/help/
    
    /admin/settings  → Blocked  (Disallow: /admin/)
    /admin/help/faq  → Allowed  (Allow: /admin/help/, longer)

    The more specific rule wins regardless of order in the file.

  • A named group replaces the * group

    User-agent: *
    Disallow: /private/
    
    User-agent: Googlebot
    Disallow: /search
    
    /private/report → Allowed for Googlebot

    Googlebot only reads its own group, so /private/ is open to it. Repeat shared rules in every named group.

  • Wildcards and end anchors

    Disallow: /*?sessionid=     blocks /cart?sessionid=abc
    Disallow: /*.pdf$           blocks /guides/pricing.pdf
                                allows /guides/pricing.pdf?v=2

    * matches anything, including slashes. $ means the URL must end there, so a query string after .pdf escapes the rule.

  • Staging site that should never be crawled

    User-agent: *
    Disallow: /

    Use it on staging only, and protect staging with a password as well. Shipping this file to production removes a site from search over time.

What robots.txt does, and what it does not

robots.txt is a plain text file at the root of a host, such as https://example.com/robots.txt, that tells crawlers which paths they may fetch. It was standardised as the Robots Exclusion Protocol in RFC 9309 in 2022, after nearly three decades as an informal convention. Each protocol, host and port needs its own file, so a subdomain does not inherit the main site's rules.

It controls crawling, not indexing. A disallowed URL can still appear in search results, without a description, if other pages link to it. To keep a page out of search, allow crawling and add a noindex robots meta tag or header, because a crawler has to fetch the page to see noindex. robots.txt is also not access control: the file is public and only well-behaved bots obey it, so protect private areas with authentication.

How crawlers choose a rule

A file is a series of groups. Each group starts with one or more User-agent lines followed by Allow and Disallow rules. A crawler uses the group that names it most specifically and ignores the rest; only if no group names it does it use the * group. That catches many people out: if you add a group for Googlebot, Googlebot no longer reads the rules under *.

Within that group, the rule with the longest matching path wins, and if an Allow and a Disallow are equally long, Allow wins. Paths are case-sensitive and match from the start. * matches any sequence of characters and $ anchors the end, so Disallow: /*.pdf$ blocks every URL ending in .pdf. The tester applies exactly these rules and shows which line decided the result.

Google ignores Crawl-delay; Bing and some other crawlers honour it. Google reads at most 500 KiB of the file, and unrecognised lines are skipped, which the status bar flags so typos do not fail silently.

AI crawlers in 2026: training, search and user fetches

AI companies now run separate agents for separate purposes, and the difference matters. Training crawlers collect content that may be used to train models: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl, whose dataset is widely used for training), Meta-ExternalAgent, Amazonbot and Bytespider. Google-Extended and Applebot-Extended are not separate crawlers but tokens: Google's regular crawlers fetch the page, and the token decides whether the content may be used for Gemini training and grounding, or for Apple's foundation models.

Search crawlers index pages so an assistant can cite and link them: OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude and PerplexityBot for Perplexity. Blocking these removes you from those answers. User-triggered fetchers, ChatGPT-User, Claude-User and Perplexity-User, load a page because a person asked about it. OpenAI and Perplexity say robots.txt may not apply to these; Anthropic says Claude-User honours it.

A common middle ground, and the default preset here, is to block training crawlers and allow search and user agents: your content is not used to train models, but you can still be cited and linked. Blocking Google-Extended does not affect Google Search rankings or crawling.

Common robots.txt mistakes

The most damaging mistake is shipping a staging file with Disallow: / to production, which asks every search engine to stop crawling the whole site. Others include blocking CSS and JavaScript files that Google needs to render pages, blocking a page you want deindexed (so Google never sees its noindex), relying on robots.txt to hide sensitive URLs (the file advertises them), and forgetting that rules are case-sensitive, so /Admin and /admin are different paths.

Crawlers cache robots.txt, typically for up to a day, so changes are not instant. If the file returns a 4xx error, crawlers treat the site as unrestricted; persistent server errors make Google pause crawling. Make sure the file returns 200 with a text/plain content type.

Sitemaps and llms.txt

Sitemap lines can appear anywhere in the file and apply to all crawlers. List the absolute URL of every sitemap or sitemap index; it is the easiest way to help search engines discover new pages on a large site. They complement, not replace, submitting sitemaps in Google Search Console and Bing Webmaster Tools.

You may also see llms.txt, a proposed Markdown file that summarises a site for language models. It is a proposal rather than a standard, and the major AI crawlers do not document using it as a permission mechanism, so it does not replace robots.txt rules. See generative engine optimization for how sites are approaching AI search more broadly.

Testing before you deploy

Paste your current live file into the text pane to test it as-is, or build a new one with the form. Then try the URLs that matter: the home page, a product or article, an admin path, a URL with tracking parameters, a PDF. The row of crawler chips shows the verdict for every bot at once, which makes it easy to spot a rule that blocks more than you intended.

After uploading, open yourdomain.com/robots.txt to confirm it is served, and check Google Search Console's robots.txt report, which shows the version Google fetched and any errors it found. For a full crawlability review of a site, our technical SEO and SEO services pages explain what else to check.

robots.txt vs meta robots vs X-Robots-Tag

These three are often confused. robots.txt works at the level of URL paths and decides whether a crawler may fetch a page. The robots meta tag, placed in a page's head, and the X-Robots-Tag HTTP header, which also works for PDFs and images, tell search engines what to do with a page they have fetched: noindex keeps it out of results, nofollow asks them not to follow its links, and nosnippet or max-snippet limit how much of it appears in results.

Use robots.txt to keep crawlers out of areas that waste crawl capacity, such as internal search results, filtered listing URLs and cart pages. Use noindex for pages that must stay reachable but should not appear in search, such as thank-you pages or thin tag archives. Never combine a robots.txt block with noindex on the same URL, because the crawler cannot fetch the page to see the noindex.

Questions, answered

Something else on your mind? Ask a consultant and get a reply within one business day.

How do I block ChatGPT from using my content?

Disallow GPTBot to opt out of OpenAI model training. OAI-SearchBot controls whether you appear in ChatGPT search; ChatGPT-User fetches pages when users ask, and OpenAI says robots.txt may not apply to it.

How do I block Claude or Anthropic?

Disallow ClaudeBot to opt out of training. Claude-SearchBot controls search indexing and Claude-User covers user-requested fetches. Anthropic says all three honour robots.txt.

What does Google-Extended do?

It is a robots.txt token, not a crawler. Disallowing it stops content Google crawls from being used to train Gemini models and for grounding in Gemini Apps and Vertex AI. It does not affect Google Search.

Will blocking AI crawlers hurt my Google rankings?

No. Googlebot is separate, and Google says Google-Extended is not a ranking signal. Blocking AI search crawlers does remove you from those assistants' answers.

Do AI crawlers obey robots.txt?

The major operators document that their training and search crawlers do. User-triggered fetchers are a grey area, and some crawlers have been reported to ignore the file. Use your CDN or firewall for hard blocking.

Does robots.txt stop a page from being indexed?

No. It stops crawling. A blocked URL can still be indexed from links. Use a noindex meta tag or header, and leave the page crawlable so search engines can see it.

Where do I put robots.txt?

At the root of each host: https://example.com/robots.txt. Subdomains such as shop.example.com need their own file.

Does Google support Crawl-delay?

No. Google ignores it and adjusts its crawl rate automatically. Bing and some other crawlers honour it.

Which rule wins when Allow and Disallow both match?

The one with the longer path. If they are the same length, Allow wins. The order of lines in the file does not matter.

Should I add llms.txt?

It is optional and still a proposal. AI crawlers do not document it as a permission mechanism, so it does not replace robots.txt rules.

Is my robots.txt sent anywhere?

No. Building, parsing and testing all happen in your browser.

More free tools.

All tools

Need tooling like this inside your product?

We build internal tools, developer platforms and APIs. Tell us what your team keeps doing by hand.