PDF Tools
PDF to text.
Pull the text out of a PDF and copy it or save it as a .txt file. The PDF is read in your browser and never uploaded. If the PDF is a scan with no real text, the page says so instead of returning nothing.
Input
Extracted text
Choose a PDF to extract its text.
About this PDF to text converter
Most PDFs created from Word, Google Docs or a website contain a real text layer: each word is stored as text with a position on the page. This tool reads that layer with the pdf.js library, groups the pieces into lines, and outputs plain text. It does not guess at pictures of text.
Scanned PDFs
A scan or a photographed document is a picture, so there is no text to extract, and the page tells you when it finds none. To get text from a scan you need OCR. A quick test: if you cannot select a single word with your cursor in a PDF viewer, it has no text layer.
What to expect in the output
- Reading order: lines are rebuilt top to bottom, left to right using each piece's position. Single-column documents come out well. Multi-column pages, sidebars and tables can interleave columns.
- Line breaks: each visual line becomes one line. For flowing paragraphs run the result through the line break remover.
- Hyphenation, headers and footers are kept exactly as they appear on the page.
- Special fonts: a PDF whose fonts map to private code points can extract as garbled characters even though text is selectable.
Limits
Up to 40 MB and 300 pages. Password-protected PDFs stop with an error. The file is read in your browser and this page loads no advertising or analytics scripts.
FAQ
How do I copy text from a PDF I cannot select?
If you cannot select any text, the PDF is an image and needs OCR. If you can select text, choose the file here, press Copy text, and paste it anywhere.
Why is the extracted text in the wrong order?
The tool follows each piece of text by its page position. Two-column layouts, tables and text boxes can interleave. Extract one page at a time with Split PDF if one page is the problem.
Is my PDF uploaded?
No. Extraction happens in your browser.
Source
- pdf.js, Mozilla's PDF rendering and text-extraction library: mozilla.github.io/pdf.js.