Extract the text from a PDF
Last reviewed: September 23, 2026 · Markdown version
Sometimes you do not want to search a PDF, you want the words out of it — to paste into a note, to diff against another version, to feed to something else. Copying from a reader can drag the layout along with the words, and running headers and page numbers too.
Choose a PDF below and this page gives you back the text it stores, one page at a time.
…or drag one here.
The file never leaves your browser. Nothing is uploaded, no network request carries it anywhere, and nothing is stored — the reading happens on your own machine, in this page. Close the tab and it is gone.
This page reads the text stored inside the PDF. It does not run OCR, so a page that is a photograph of words comes back marked as having no text layer rather than as an empty page — the difference matters, and no tool should hide it from you.
What you get, and what you do not
The words, in the order the file stores them
A PDF does not store sentences. It stores instructions to draw pieces of text at positions on a page — in the standard's words, "operators that may show text strings, move the text position, and set text state" — and the order they are drawn in need not be the order you read them. The standard says a page's drawing order and its logical reading order "may or may not coincide" (PDF 32000-1:2008, Adobe's published copy of the PDF standard, §9.4.1 and §14.8.2.3.1, read 23 September 2026). Text in two columns, in a table, or in a sidebar can come back interleaved. What you see here is the same raw material a PDF text search works from, which is exactly why a search sometimes behaves in ways that look strange.
Nothing from a scanned page
If the page is an image, there is no text in the file to extract — not for this page, and not for any tool that only reads the text a PDF stores. Those pages are marked. Turning them into text means running OCR over them first: what OCR is, and how to make a scanned PDF searchable.
Files that need a password to open
The standard lets a PDF carry a user password, needed to open it, and an owner password, which governs permissions such as printing and copying (same standard, §7.6.3.1). If a file needs a user password to open, this page cannot read it. If it has only an owner password, it opens without asking and its text comes out here normally — and DocFind reads it as it would any unprotected PDF.
The honest limits
- This page does not run OCR. Scanned pages come back empty and labelled.
- Layout is not preserved. Columns, tables and headers come out as text in drawing order, not as a reproduction of the page.
- Words hyphenated across a line break stay in two pieces, because that is how the file stores them.
- Nothing is remembered. Reload the page and it is empty again.
- PDFs only. Word, Excel and plain-text files cannot be read here.
Looking for a word rather than the whole thing?
Search across many PDFs at once takes a folder and gives you the file, the page and the sentence around every match. And DocFind does that over a whole library on your phone, reading scanned pages with OCR on the device — nothing uploaded, the index built and kept locally. iPhone and Android.