# Extract the text from a PDF

*Last reviewed: September 23, 2026 · [HTML version](https://docfindapp.com/tools/extract-text-from-pdf)*

Sometimes you do not want to search a PDF, you want the words out of it — to paste into a note, to
diff against another version, to feed to something else. Copying from a reader can drag the layout along
with the words, and running headers and page numbers too.

The tool on the HTML version of this page takes a PDF and gives back the text it stores, one page at
a time, ready to copy or to save as a `.txt` file. **The file never leaves your browser** — nothing
is uploaded, no network request carries it anywhere, and nothing is stored.

## What you get, and what you do not

**The words, in the order the file stores them.** A PDF does not store sentences. It stores
instructions to draw pieces of text at positions on a page — in the standard's words, "operators that
may show text strings, move the text position, and set text state" — and the order they are drawn in
need not be the order you read them. The standard says a page's drawing order and its logical reading
order "may or may not coincide" (*PDF 32000-1:2008*, Adobe's published copy of the PDF standard, §9.4.1
and §14.8.2.3.1, read 23 September 2026 —
https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandards/PDF32000_2008.pdf). Text in two columns,
in a table, or in a sidebar can come back interleaved. What you see is the same raw material a PDF text
search works from, which is exactly why a search sometimes behaves in ways that look strange.

**Nothing from a scanned page.** If the page is an image, there is no text in the file to extract —
not for this page, and not for any tool that only reads the text a PDF stores. Those pages are marked
rather than returned as empty. Turning them into text means running OCR over them first:
[what OCR is](https://docfindapp.com/answers/what-is-ocr), and
[how to make a scanned PDF searchable](https://docfindapp.com/answers/make-a-scanned-pdf-searchable).

**Files that need a password to open.** The standard lets a PDF carry a **user password**, needed to
open it, and an **owner password**, which governs permissions such as printing and copying (same
standard, §7.6.3.1). If a file needs a user password to open, this page cannot read it. If it has only
an owner password, it opens without asking and its text comes out here normally — and **DocFind reads
it as it would any unprotected PDF**.

## The honest limits

- **This page does not run OCR.** Scanned pages come back empty and labelled.
- **Layout is not preserved.** Columns, tables and headers come out as text in drawing order, not as
  a reproduction of the page.
- **Words hyphenated across a line break** stay in two pieces, because that is how the file stores
  them.
- **Nothing is remembered.** Reload the page and it is empty again.
- **PDFs only.** Word, Excel and plain-text files cannot be read here.

## Looking for a word rather than the whole thing?

[Search across many PDFs at once](https://docfindapp.com/tools/search-multiple-pdfs) takes a folder
and gives you the file, the page and the sentence around every match. And DocFind does that over a
whole library on your phone, reading scanned pages with OCR *on the device* — nothing uploaded, the
index built and kept locally. iPhone and Android: https://docfindapp.com/#get

Related: [other free tools](https://docfindapp.com/tools/) ·
[all answers](https://docfindapp.com/answers/)
