Extract clean text from a website, without the clutter
Copying a web page gives you menus, cookie notices, share buttons and broken formatting. Maybe you need the text of an article, a policy page or a whole docs section to read, quote or translate. Your todo.is agent opens the pages in a real browser, keeps only the main content with its headings and lists, and hands you a tidy file.
The prompt
- Extract the main text from [WEBSITE URLS]. Keep only the content: headings, paragraphs, lists, tables and image captions. Remove menus, ads, cookie banners, comments, related-article boxes and footers. [WHICH PAGES]. Keep links as [LINK FORMAT]. At the top of each page write its title, URL and the date you copied it. Give me a [FILE FORMAT]. If a page needs a login or blocks access, tell me instead of guessing.
What to change
- [WEBSITE URLS]: One or more links, e.g. "https://example.com/blog/guide-to-composting". Public pages only.
- [WHICH PAGES]: E.g. "just this page", "also follow the Next links to get all 5 parts" or "every page listed in the left menu of the docs".
- [LINK FORMAT]: E.g. "clickable links", "footnotes with the URL at the end" or "plain text, no links".
- [FILE FORMAT]: E.g. "Word document", "Markdown file", "plain .txt" or "one PDF per page".
Example result
- Extract report
- Source: a 5-part composting guide on a gardening blog (public pages)
- Result: Composting_guide_text.docx, 5 sections, 7,840 words, 3 tables
- What was kept
- • The article title and every H2 and H3 heading, as real Word headings (so the navigation pane works)
- • Paragraphs, numbered steps and bullet lists
- • The "Greens vs browns" table, rebuilt as a Word table
- • Image captions, marked as [Image: ...] where the picture was
- • 42 links, as clickable text
- What was removed
- • Top menu, search box and newsletter pop-up
- • Cookie banner and ads between paragraphs
- • "You might also like" boxes and the comment section
- • Share buttons and the footer
- Sample of the result
- Part 2: Getting the balance right
- Source: https://example-garden-blog.com/composting/part-2 · copied 8 Oct 2026
- A healthy pile needs both greens and browns. Greens are fresh, wet materials rich in nitrogen. Browns are dry, woody materials rich in carbon.
- • Greens: vegetable scraps, coffee grounds, grass clippings
- • Browns: dry leaves, cardboard, straw
- [Image: A cross-section of a compost bin showing layers]
- Good to know
- • Text inside images isn't copied by default. Ask for OCR if a page shows text as pictures
- • Pages that load content as you scroll are read in a real browser, so the full text is captured
- • Pages behind a login can be read through your own Chrome with the todo.is Connector, only for accounts you own
- Next steps you could ask for
- • "Summarize each part in 5 bullets"
- • "Translate the whole thing into German, same headings"
How to do it with todo.is
- Copy the prompt and paste in the links. Say whether to follow Next pages or a menu.
- Send it from todo.is or to your agent on WhatsApp, Telegram or email.
- Your agent opens each page, strips the clutter and builds one clean file.
- Download it, or ask for a different format, a summary or a translation of the text.
Tips for a better result
- For a multi-page article or a docs site, say exactly how far to go ("all pages under /docs/setup"), so you don't get the whole website.
- Pick Markdown if you'll paste the text into another tool or an AI chat. Pick Word if people will read or edit it.
- Keep the source URL and date at the top. Web pages change, and you'll want to know where a quote came from.
- For data in rows and columns (prices, directories), a table extract to Excel is a better fit than plain text.
- Respect copyright: extracting text for reading, research or notes is fine. Republishing someone else's article is not.
extract text from a website: FAQ
- How do I copy text from a website that blocks copying? The text is still in the page, so your agent can read it in the browser and save it. Use it for personal reading or research, and respect the site's terms.
- Can I extract text from many pages at once? Yes. Give a list of links, or one starting page and the rule to follow (Next buttons, a menu, or a URL path). You get one file with a section per page.
- Can it extract text from images on a website? Yes, with OCR. Ask for it, and your agent reads the text in the pictures and adds it under each image.
- Can it get text from a page behind a login? Only for accounts you own, through your own Chrome with the todo.is Connector. Your agent never asks for or stores your password.
JavaScript is required to use the todo.is app.