PDF to Text
Extract a PDF's text into a plain .txt file, page by page.
Need help? Read our guide on extracting text from a PDF.
Extract Text from a PDF
This tool reads the actual text layer embedded in your PDF — the same underlying text you could select and copy in a PDF viewer — and writes it out to a plain .txt file, one page's text after another, separated by a blank line. It doesn't try to preserve columns, tables, or formatting; the output is deliberately plain so it's easy to search, paste into another document, or feed into another tool.
This only works on PDFs that already contain real text. If your PDF is a scan or a photograph of a document, there's no text layer to read — those pages have no embedded text at all, only pixels — and this tool has no OCR (optical character recognition) to read text out of an image. Pages with no extractable text are simply skipped, which can mean an empty or near-empty result for a fully scanned document.
How it works
- 1 Upload your PDF
- 2 The text layer of every page is extracted
- 3 Download a .txt file with the same name as your PDF
When to use this
- Reusing content: pull the wording out of a report, contract, or article without retyping it.
- Searching a large document: a plain text file is easy to search with any text editor once it's out of the PDF.
- Feeding text into another tool: word counters, translators, or anything that needs raw text rather than a PDF.
- Quick content checks: confirm what text is actually embedded in a PDF you didn't create yourself.
Frequently Asked Questions
Will this work on a scanned document?
Not reliably. A scanned PDF is really a sequence of page images with no embedded text, so there's nothing for this tool to extract — it doesn't perform OCR. If you need text out of a scanned document, you'll need an OCR-specific tool.
Does it preserve tables, columns, or formatting?
No. The output is plain text with each page's content separated by a blank line — column layouts, tables, and font styling are not preserved. If you need to keep formatting, use PDF to Word instead.
Why is my output file empty or missing some pages?
Pages with no extractable text (usually scanned pages, or pages that are entirely images) are skipped rather than producing garbled output — an empty result means the source PDF had no embedded text at all.
Is the text extracted in reading order?
Text is pulled out in the order it's stored in the PDF, which usually matches reading order for simple single-column documents. Multi-column layouts, sidebars, or complex page designs can come out in a different order than you'd read them visually, since the tool doesn't try to detect columns.