Where can I find research papers — and how do I search the ones I have?
Last reviewed: September 23, 2026 · Markdown version
A great deal of the literature is legitimately free, and the awkward part comes
afterwards — when 300 PDFs with names like 1706.03762.pdf are sitting in a folder and
you cannot remember which one had the method.
Where the free copies actually are
- PubMed Central for biomedical and life sciences — "a free full-text archive of biomedical and life sciences journal literature at the U.S. National Institutes of Health's National Library of Medicine", which "makes all content free to read (in some cases, following an embargo period)" (About PMC, read 23 September 2026).
- arXiv for physics, mathematics, computer science and the other fields it lists — "a curated research-sharing platform open to anyone", with "no fees or costs for article submission." Read its own caveat before citing anything: "Material is not peer-reviewed by arXiv - the contents of arXiv submissions are wholly the responsibility of the submitter" (About arXiv, read 23 September 2026). A preprint is a draft in public, not a published finding.
- Institutional repositories. Many universities run one for their own researchers' work, and it is worth checking the author's institution for an accepted manuscript.
- The author. A short, polite email asking for a copy costs nothing to send.
Publisher paywalls are a real constraint and there are honest ways around them — your library, interlibrary loan, an author copy. Sites that host pirated copies are a different thing, and this page is not going to point you at them.
The problem that starts once you have them
A downloaded paper is filenamed by whoever exported it, which means an arXiv id, a DOI fragment,
or sdarticle(3).pdf. Six months later you remember a phrase from a methods section and
have no idea which file it was in. Folder names and tags do not help, because what you remember is
content, not category.
Searching the text is the only reliable route back — and it has to search every paper at once, not one open document at a time.
Two wrinkles specific to papers
- Old scans have no text layer. Older and historical material that was digitised as page images is a picture of a page. It needs OCR before it is searchable at all.
- Two-column layouts complicate extraction. Pulling text out means a tool has
to decide the reading order. The standard
pdftotexttool, for instance, by default undoes the physical layout — "(columns, hyphenation, etc.)" — to "output the text in reading order", with a separate option to "Maintain (as best as possible) the original physical layout" (pdftotext(1) manual page, read 23 September 2026) — one more reason to search inside the document rather than extract and grep.
Where DocFind fits
DocFind is for the second half of this page. Point it at your papers folder and it indexes the lot, then searches inside all of them at once — results show the file, the page and the surrounding text, and tapping one opens that paper at that page with your terms highlighted. Scanned material is read on the device first, so old PDFs join the searchable set without going anywhere near a server. It runs entirely offline, and your documents never leave the phone.
It does not fetch papers, manage citations, or talk to a reference manager. It finds the sentence you half-remember, in the file you cannot name.
Related: Search many PDFs at once · Search for a word in a PDF · Make a scanned PDF searchable · Searching without uploading