PDF to text converter that keeps reading order and fixes broken lines
Copying text out of a PDF gives you broken lines, words split by hyphens, page headers in the middle of sentences and columns mixed together. Your todo.is agent extracts the text in the right reading order, joins lines back into paragraphs and runs OCR on scanned pages, so you get text you can actually use.
The prompt
- Extract the text from the attached PDF: [ATTACH THE PDF]. I need [PAGES]. Keep the reading order, including for [LAYOUT]. Join lines back into paragraphs, re-join words split by hyphens at line ends, and remove page headers, footers and page numbers. Run OCR on any scanned pages. Mark each page with "--- Page N ---". Keep tables as [TABLE FORMAT]. Send me a UTF-8 .txt file and tell me which pages needed OCR.
What to change
- [ATTACH THE PDF]: Attach the PDF, up to 50 MB. Text-based and scanned PDFs both work.
- [PAGES]: E.g. "all pages", "pages 12 to 40" or "only chapter 3".
- [LAYOUT]: E.g. "two-column pages", "footnotes", "sidebars" or "nothing special".
- [TABLE FORMAT]: E.g. "tab-separated rows", "Markdown tables" or "skip tables".
Example result
- PDF to text: Annual_Report_2025.pdf
- 64 pages · 58 text pages + 6 scanned pages (OCR) · 31,400 words · UTF-8 .txt
- Clean-up done
- • Reading order: pages 8 to 40 are in two columns. Text runs down the left column, then the right, not across both
- • Line breaks: lines were joined into full paragraphs; real paragraph breaks kept
- • Hyphenation: "manage-\nment" became "management" (1,214 fixes)
- • Removed: the running header "Northwind Annual Report 2025", footers and page numbers
- • Footnotes: moved to the end of their page, marked [1], [2]
- • Ligatures and special characters: "fi" and "fl" became "fi" and "fl", curly quotes kept
- Example of the result
- --- Page 12 ---
- Chairman's letter
- This year we opened two new distribution centers and reduced delivery times across the region. The board approved a plan to invest in electric vans over the next three years.[1]
- [1] See the sustainability section on page 41.
- Tables
- Tables were kept as tab-separated rows, so you can paste them into Excel and they split into columns.
- OCR pages
- Pages 1, 2, 61, 62, 63 and 64 were scans. They were read with OCR. Check numbers on page 62, which is slightly blurry.
- Want a different format?
- Ask for Markdown with headings, a Word file, or just the text of one section.
How to do it with todo.is
- Copy the prompt, choose the pages and how tables should look.
- Attach the PDF in todo.is, or send it to your agent on Telegram or WhatsApp.
- Your agent extracts the text, fixes the layout problems and runs OCR where needed.
- Download the .txt. Ask for a summary, a translation or Markdown from the same text next.
Tips for a better result
- Mention two-column layouts. Newspapers, papers and reports often use them and they are the main cause of jumbled text.
- If you plan to feed the text into another AI tool or search index, ask for page markers. You can trace answers back to a page.
- For scans, a clear 300 dpi file gives far better OCR than a phone photo of a page.
- Need only facts from the text? Ask your agent to extract them into a table instead of a full text dump.
PDF to text converter: FAQ
- Can I extract text from a scanned PDF? Yes. Your agent runs OCR on scanned pages and tells you which pages were read that way, so you know which ones to double-check.
- Why is copied PDF text full of line breaks? PDFs store text as positioned lines, not paragraphs. Your agent joins the lines back into paragraphs and removes hyphens at line ends.
- Can it extract text from a password-protected PDF? Only if you give the password and the file is yours. Your agent does not remove protection from files you do not own.
- Will it keep headings and bold text? A .txt file holds only plain text. Ask for Markdown or Word if you want headings and formatting kept.
JavaScript is required to use the todo.is app.