Website to Excel scraper that works from a plain-English request
You found a directory, catalog or listings page with exactly the data you need, but copying it row by row would take all afternoon. This prompt has your todo.is agent open the site in a real browser, follow the pages, pull out the fields you name and hand you a tidy Excel file with one row per item and the source link for each.
The prompt
- Scrape [WEBSITE OR PAGE LINK] into an Excel file. For each item on the page, collect these columns: [COLUMNS YOU WANT]. Follow the "next page" links and stop after [HOW MANY PAGES OR ITEMS]. Open each item's detail page only if a column is not on the list page. Add a "Source URL" column for every row. Clean the data: trim spaces, keep prices as numbers with a separate currency column, use one date format (YYYY-MM-DD) and leave a cell empty rather than guessing. Remove exact duplicates. Save it as [FILE NAME].xlsx with a frozen header row and filters on, plus a second sheet listing any pages you could not read and why. Only use public pages and respect the site's terms and robots.txt.
What to change
- [WEBSITE OR PAGE LINK]: The public page where the list starts, e.g. a category page, a directory search result or an events listing.
- [COLUMNS YOU WANT]: e.g. "name, price, rating, number of reviews, brand" or "company, city, phone, website".
- [HOW MANY PAGES OR ITEMS]: e.g. "10 pages" or "300 items". Start small to check the result first.
- [FILE NAME]: e.g. "Coffee_Grinders_Oct". Your agent adds the .xlsx.
Example result
- Coffee_Grinders_Oct.xlsx
- Source: public category page "Coffee grinders" on an example kitchen store · 12 pages read · 286 rows
- Sheet 1: Data (first rows)
- • Name: Burrline Pro 40 · Brand: Burrline · Price: 189.00 · Currency: EUR · Rating: 4.6 · Reviews: 412 · Source URL: /p/burrline-pro-40
- • Name: Mokka Hand Grinder · Brand: Mokka · Price: 59.90 · Currency: EUR · Rating: 4.3 · Reviews: 88 · Source URL: /p/mokka-hand
- • Name: Steelmill S2 · Brand: Steelmill · Price: 129.00 · Currency: EUR · Rating: (empty, no reviews yet) · Reviews: 0 · Source URL: /p/steelmill-s2
- How the file is set up
- • Header row frozen, filters on every column
- • Price and rating stored as numbers, so you can sort and use SUM or AVERAGE
- • Currency in its own column instead of mixed into the price text
- • 9 exact duplicates removed (the same product shown in two places)
- Sheet 2: Not collected
- • Page 7, item "Grindmaster Duo": detail page returned an error, retried once
- • Page 11, item "Bundle offer": no price shown on the page, left empty
- Notes
- • The site's robots.txt allows category pages. I paused between pages so the site was not overloaded.
- • Prices are what the page showed on Oct 8. Sale prices were taken when shown, with the regular price in a separate column.
- Want this every week? Say "run this every Monday at 7:00 and send me only new or changed rows".
How to do it with todo.is
- Copy the prompt and fill in the page link, the columns and how much to collect.
- Paste it on the Today screen in todo.is, or send it to your agent on WhatsApp or Telegram.
- Your agent browses the pages, collects the fields and checks the data before building the file.
- Download the .xlsx, look over a few rows against the site, then ask for extra columns or a weekly rerun.
Tips for a better result
- Name your columns exactly as you want them in the header. It saves a rename later.
- Run a small test first (one or two pages), check it, then ask for the full run.
- For pages you must log in to see, connect your own Chrome with the todo.is Connector, and only collect data you are allowed to use.
- Ask for a "Scraped on" date column if you will compare runs over time.
- If a site offers an official export or API, use it. It is more reliable than scraping.
website to Excel scraper: FAQ
- Is it legal to scrape a website into Excel? Collecting public data for your own use is common, but rules depend on the site's terms, your country and what the data is. Avoid personal data and check the terms before reusing or publishing it.
- Can it scrape pages that load with JavaScript or infinite scroll? Yes. Your agent uses a real browser, so it can scroll and wait for content to load like a visitor would.
- How many rows can I get? Thousands is fine for most sites, but bigger runs take longer and use more credits. Start with a few pages and scale up.
- Can I get CSV or Google Sheets instead of Excel? Yes. Ask for CSV, or ask your agent to put it in Google Sheets if your Google account is connected.
JavaScript is required to use the todo.is app.