Robots.txt generator

Start from a template, adjust the rules for each crawler, declare your sitemap. The file updates live, ready to copy or download.

Free, no sign-up

Starting template

A template replaces the groups below. You can edit them afterwards.

AI crawlers

AI assistants can read your pages, use them for training and cite them in their answers.

Rules per crawler

Group 1
One crawler per line. The asterisk means every crawler that has no group of its own.
Full address, one sitemap per line.

robots.txt

User-agent: *
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-admin/
Disallow: /?s=
Disallow: /search/

    Upload this file to the root of your domain: it must answer at the /robots.txt address.

    What robots.txt is for

    robots.txt is a text file at the root of your domain, for example https://www.your-site.com/robots.txt. It tells crawlers which parts of the site they may crawl and which they should leave alone. Each group of rules starts with one or more User-agent lines, naming the crawlers concerned, followed by Disallow and Allow lines.

    The key point: robots.txt controls crawling, not indexing. A blocked page can still appear in Google if other pages link to it, just without a description. To remove a page from the results you need a meta robots noindex tag, and the page must stay crawlable so the crawler can read that tag.

    How to use the generator

    • Pick a starting template: allow everything, WordPress (which closes the admin area and internal search pages) or block everything for a staging site.
    • Decide where you stand on AI crawlers. The generator adds the matching group of rules for you.
    • Adjust the groups: one path per line, starting with a slash. The asterisk is a wildcard, the dollar sign marks the end of a URL.
    • Add the full address of your XML sitemap, then copy the file or download it and upload it to the root of the site.

    AI crawlers: training or answers

    Not all AI crawlers do the same thing. Some collect pages to train models: GPTBot at OpenAI, ClaudeBot at Anthropic, Google-Extended for the Gemini models, Applebot-Extended, CCBot for Common Crawl. Others fetch a page at the moment someone asks a question, to answer it and cite the source: OAI-SearchBot and ChatGPT-User, Claude-SearchBot and Claude-User, PerplexityBot and Perplexity-User.

    Blocking training crawlers has no effect on your Google rankings: Google-Extended does not concern search. Blocking answer crawlers, on the other hand, means disappearing from those assistants' answers, since they can no longer read or cite your pages. If your visibility in AI answers matters, the middle option is often the right one.

    Common mistakes

    • Leaving the staging Disallow: / in place after launch: the whole site gradually drops out of Google.
    • Blocking CSS and JavaScript files, which Google needs to render the page the way a visitor sees it.
    • Blocking a page to deindex it: the crawler can no longer see the noindex tag and the URL may stay in the results.
    • Forgetting that paths are case-sensitive: /Blog/ and /blog/ are two different rules.
    • Putting the file in a subfolder. It must sit at the root, and each subdomain needs its own.

    These tools look at one page. Bloomwise looks at your whole site.

    14-day trial · No card required · Cancel anytime

    Add your site: Bloomwise audits all of it, gives you an SEO score and the list of fixes to make, then helps you produce the content that is missing.