DocFind
Features Pricing Support Answers Tools Get DocFind

How do I tell whether a PDF has real text in it?

Last reviewed: September 23, 2026 · Markdown version

Try to select a word. Open the PDF, drag across a line of text or long-press it on a phone. If individual words highlight, the file contains real text. If the whole page turns into one blue rectangle — or nothing happens at all — the page is an image, and no amount of searching will find a word that is not stored anywhere.

Two more tests, when selection is ambiguous

  • Search for a word you can see. Pick something unmistakable on screen and use the reader's find function. A text-layer PDF finds it; an image PDF reports nothing, which is the single most common "PDF search is broken" complaint. The longer diagnosis.
  • Look at the file size per page. A hundred text pages are often smaller than a handful of scanned ones. Megabytes per page is a strong hint that you are holding photographs.

The confusing middle case

A PDF can be partly searchable, and this catches people out. A report typed on a computer with three scanned appendices bolted on has a text layer for some pages and none for the rest, so search works until it silently doesn't. The same happens when OCR has been run once, badly: some pages carry recognised words, others were skipped, and nothing on screen tells you which.

A scanned page that has been through OCR also keeps looking exactly like a photograph — because it still is one. The open-source OCRmyPDF project describes its whole job as adding text "layers" to the images in a PDF, "making scanned image PDFs searchable" (OCRmyPDF: Introduction, read 23 September 2026): the picture stays, and the words sit with it where you cannot see them. That is why the selection test, not your eyes, is the one that answers the question.

What to do with each answer

  • It has text — copying and searching work now, in the reader you already have, at no cost.
  • It does not — it needs OCR before any of that is possible. How to run it, and where it runs matters: an online converter only works once the document has been uploaded to it, which is a poor trade for anything confidential.

Where DocFind fits

DocFind answers this question for a whole library instead of one file at a time. It indexes your PDFs, reads scanned ones on the device — Vision on iPhone, a bundled ML Kit model on Android, both working in airplane mode — and then tells you what it could not do: which files were skipped and why, and which pages are not searchable yet. Matches that came from OCR are labelled as such, so you always know whether a result came from the file's own text or from recognition that might have misread a word.

Related: What is OCR? · Make a scanned PDF searchable · PDF search not finding words? · Extracting text from a PDF

Know someone with this problem? Share this answer
WhatsApp X Facebook Reddit Telegram Email

These are ordinary links — nothing is loaded and nobody is told anything until you pick one.

DocFind
Features Pricing Privacy Policy Support Answers Tools

Questions or bugs? support@docfindapp.com
DocFind is not affiliated with, endorsed, or sponsored by Apple Inc. or Google LLC. Apple, the Apple logo and iPhone are trademarks of Apple Inc. Google Play and the Google Play logo are trademarks of Google LLC.