Robots.txt generator that explains every line
One wrong slash in robots.txt can hide your whole site from Google, and many files copied from templates block CSS or images by mistake. With this prompt your todo.is agent reads your current file and site structure, writes a clean robots.txt for your goals, explains each rule and tests your important URLs against it.
The prompt
- Write a robots.txt file for [WEBSITE]. First read my current robots.txt and sitemap if they exist. I want to block crawling of [WHAT TO BLOCK]. For AI crawlers, I want to [AI CRAWLER CHOICE]. Make sure CSS, JavaScript, images and my important pages stay crawlable. Add the Sitemap line. Then test these URLs against the new rules and say for each whether it is allowed or blocked: [URLS TO TEST]. Explain every line in plain words, point out anything risky in my current file, and remind me which pages need a noindex tag instead of a block. Give me the final file as robots.txt.
What to change
- [WEBSITE]: e.g. "https://shop.moorlandtea.co.uk".
- [WHAT TO BLOCK]: e.g. "cart, checkout, account pages, internal search results and filter URLs with ?sort=".
- [AI CRAWLER CHOICE]: e.g. "allow them", "block AI training crawlers like GPTBot and CCBot" or "decide for me and explain".
- [URLS TO TEST]: 5 to 10 URLs, e.g. your homepage, a product page, a search URL and a CSS file.
Example result
- Your new robots.txt (example for shop.moorlandtea.co.uk)
- User-agent: *
- Disallow: /cart
- Disallow: /checkout
- Disallow: /account
- Disallow: /search
- Disallow: /*?sort=
- Disallow: /*?filter=
- User-agent: GPTBot
- Disallow: /
- User-agent: CCBot
- Disallow: /
- Sitemap: https://shop.moorlandtea.co.uk/sitemap.xml
- What each part does
- • User-agent: \* applies the rules below to all crawlers
- • Disallow: /cart, /checkout, /account: private, no search value
- • Disallow: /search: internal search results create endless thin pages
- • Disallow: /\*?sort= and /\*?filter=: stops crawlers wasting time on the same product list in many orders. The * matches any path
- • GPTBot and CCBot blocked: these crawlers collect content for AI training. Blocking them does not affect Google Search
- • Sitemap: tells crawlers where your page list is
- Problems in your old file
- • Disallow: /assets/ blocked your CSS and images, so Google could not render pages properly. Removed
- • Crawl-delay: 10 is ignored by Google. Removed to avoid confusion
- URL test
- • / → allowed
- • /products/earl-grey-loose → allowed
- • /search?q=green → blocked
- • /collections/black-tea?sort=price → blocked
- • /assets/theme.css → allowed
- • /checkout → blocked
- Remember
- robots.txt stops crawling, not indexing. A blocked URL can still appear in Google if other sites link to it. For pages that must stay out of results, use a noindex tag and keep them crawlable.
How to do it with todo.is
- Copy the prompt, add your site and what you want blocked.
- Paste it into todo.is, or send it to your agent on Telegram or by email.
- Your agent reads your current file, writes the new one and tests your URLs.
- Upload robots.txt to the root of your domain and check it in Search Console's robots.txt report.
Tips for a better result
- Never use "Disallow: /" under User-agent: * unless you really want to hide the whole site, e.g. a staging copy.
- Do not block CSS, JS or image folders. Google needs them to see your pages as visitors do.
- Use noindex, not robots.txt, for thank-you pages or thin pages you want out of search.
- On Shopify and some hosted platforms you edit robots.txt through a template. Tell your agent the platform for exact steps.
robots.txt generator: FAQ
- Where does robots.txt go? In the root of your domain, at https://yourdomain.com/robots.txt. Each subdomain needs its own file.
- Does robots.txt stop a page from showing in Google? Not reliably. It stops crawling, but a URL can still be indexed from links. Use a noindex tag to keep a page out of results.
- How do I block AI crawlers? Add a User-agent group for each crawler, such as GPTBot or CCBot, with "Disallow: /". Well-behaved crawlers follow it, but it is a request, not a lock.
- Is robots.txt case sensitive? Paths are case sensitive, so /Cart and /cart are different. Directive names like Disallow are not.
JavaScript is required to use the todo.is app.