# How do I extract text from a PDF? — DocFind

Last reviewed: September 23, 2026. HTML version: https://docfindapp.com/answers/extract-text-from-a-pdf

**First find out whether your PDF already contains text.** Open it and try to select a sentence with your cursor or finger. If the text highlights, the words are in the file and getting them out is trivial. If you get a blue box over the whole page, the page is a picture of text and nothing can copy it until OCR runs.

## If the text selects: copy, or export

Select and copy works in ordinary PDF readers. For a whole document, many desktop readers can save or export the text, and on a computer the free command-line tool `pdftotext` "converts Portable Document Format (PDF) files to plain text" — one file per run, so a folder takes a small shell loop ([*pdftotext(1) manual page*](https://manpages.debian.org/testing/poppler-utils/pdftotext.1.en.html), read 23 September 2026). Nothing needs uploading.

## If the text does not select: OCR first

A scanned page is a photograph. Optical character recognition looks at the image, recognises the letters, and stores them as text alongside the page — after which copying, searching and exporting all work normally. [What OCR is and where it fails](https://docfindapp.com/answers/what-is-ocr), and [how to run it](https://docfindapp.com/answers/make-a-scanned-pdf-searchable).

The privacy question is worth pausing on: an online OCR site can only read the document once you have uploaded it. For a lecture handout, fine. For a contract, a payslip or a medical letter, you have put a copy on someone else's infrastructure to save yourself a tap. Phone apps have another option. Apple's Vision framework "provides pretrained machine learning models" to iPhone apps, including for "Recognizing text in 26 languages across everyday objects, documents, and photos" ([*Vision*](https://developer.apple.com/documentation/vision), read 23 September 2026); on Android, Google states that "ML Kit’s processing happens on-device" ([*ML Kit*](https://developers.google.com/ml-kit), read 23 September 2026). So for an app built on either, uploading is a choice rather than a necessity.

## "Extract to Word/Excel" is a different job

Converting a PDF into an editable document is conversion, not extraction, and it is the request most likely to disappoint. Google is candid about it for its own converter: when Drive turns a PDF into a Google Doc, "Lists, tables, columns, footnotes, and endnotes are not likely to be detected" ([*Convert PDF and photo files to text — Google Drive Help*](https://support.google.com/drive/answer/176692), read 23 September 2026). If conversion is genuinely what you need, use a converter built for it and expect to fix the layout by hand. Be especially wary of free online converters for anything confidential, for the reason above.

## The question underneath, usually

Most people asking how to extract text do not want the text. They want *one fact* out of a pile of PDFs — a policy number, a date, a clause — and extraction is the only route they can think of. Searching the pile directly skips the step entirely. [How to search many PDFs at once.](https://docfindapp.com/answers/search-inside-multiple-pdfs)

Related: [What is OCR?](https://docfindapp.com/answers/what-is-ocr) · [Make a scanned PDF searchable](https://docfindapp.com/answers/make-a-scanned-pdf-searchable) · [PDF search not finding words?](https://docfindapp.com/answers/pdf-search-not-working) · [Search many PDFs at once](https://docfindapp.com/answers/search-inside-multiple-pdfs)
