DocFind
Features Pricing Support Answers Tools Get DocFind

What is full-text search in documents?

Last reviewed: October 3, 2026 · Markdown version

Full-text search finds documents by the words inside them, not by the names they were saved under. Ask it for "indemnity" and it returns every file whose text contains that word, including the one called scan_0042.pdf. Filename search, which is what most file browsers do first, only reads the label on the outside. Below: how the people who build search engines define the term, one Windows setting that shows the difference in a single switch, and the two things that decide what "found" actually means.

The definition, from the people who build it

SQLite's documentation for its search module puts it plainly: "In their most elementary form, full-text search engines allow the user to efficiently search a large collection of documents for the subset that contain one or more instances of a search term." It adds that Google's web search "is, among other things, a full-text search engine" (SQLite FTS5 Extension, read 3 October 2026). PostgreSQL's manual adds ranking to the idea: full-text searching "provides the capability to identify natural-language documents that satisfy a query, and optionally to sort them by relevance to the query" (PostgreSQL Documentation: 12.1. Introduction, read 3 October 2026). So two parts: find the documents that contain the words, then, optionally, put the best ones first.

Filename search versus full-text search, in one Windows setting

Windows makes the distinction visible. Microsoft's indexing page says there are "two options for how much of a file to index: either properties only, or properties and content." With properties only, "indexing will not look at the contents of the file or make the contents searchable. You'll still be able to search by file name—just not file contents." By default, it says, file names and paths are indexed, and "For files with text, their contents are indexed to allow you to search for words within the files" (Search indexing in Windows, Microsoft Support, read 3 October 2026). On a phone the built-in file search is usually the first kind; what Apple's Files search covers is a worked example.

Why it needs an index

You could search contents by opening every file and reading it. PostgreSQL lists why that approach falls short: plain pattern operators "tend to be slow because there is no index support, so they must process all documents for every search." The alternative is that "Full text indexing allows documents to be preprocessed and an index saved for later rapid searching" (the PostgreSQL page above). Microsoft describes the cost side for Windows: the first indexing run "can take up to a couple hours to complete", and as a rule of thumb "the index will be less than 10 percent of the size of the indexed files" (the Microsoft page above). Those are Windows numbers, not a law of nature, but the shape holds everywhere: a one-off read of everything, then quick lookups. How an index trades space for speed goes further.

Two things that decide what "found" means

Word forms. Engines differ on whether "satisfy" should find "satisfies". PostgreSQL's manual treats missing that as a weakness of plain matching, warning that "You might miss documents that contain satisfies, although you probably would like to find them when searching for satisfy", and its preprocessing "often involves removal of suffixes (such as s or es in English)" (the PostgreSQL page above). Other tools match the word as you typed it. Neither choice is wrong, but you should know which one you are using before you conclude a word is absent.

Text that is not really text. Full-text search reads the stored text of a file. A scanned page is a photograph of words with nothing stored behind it, so it is invisible to the search until text recognition has run on it. What OCR does, and how to tell when a PDF has no text behind the picture.

Where DocFind fits

DocFind is full-text search for the PDFs on your phone: it searches inside many PDFs at once, and results show the file, the page and the surrounding text. Scanned PDFs are read on the device first and then searched like any other document. Its matching is the plain kind: there is no stemming, so searching for "running" won't find "run". On Android a completed search matches the beginnings of words, which means "run" would also list "running". It indexes PDFs only, not Word or Excel files, and on Android its recognition of scanned pages covers printed Latin-script text.

Related: Search inside multiple PDFs · Search for a word in a PDF · What is OCR? · Search inside PDFs without opening them

Know someone with this problem? Share this answer
WhatsApp X Facebook Reddit Telegram Email

These are ordinary links — nothing is loaded and nobody is told anything until you pick one.

DocFind
Features Pricing Privacy Policy Support Answers Tools

Questions or bugs? support@docfindapp.com
DocFind is not affiliated with, endorsed, or sponsored by Apple Inc. or Google LLC. Apple, the Apple logo and iPhone are trademarks of Apple Inc. Google Play and the Google Play logo are trademarks of Google LLC.