Make a scanned PDF searchable, without changing how it looks
By the getPDF team · Published 11 October 2026
The short answer
A scanned PDF is a stack of pictures of pages, so search and select find nothing in it. OCR fixes that: it reads each page image on your device and puts every recognised word back into the file as an invisible text layer, placed exactly over the printed word. The page looks identical, but it becomes selectable, searchable and copyable in any PDF reader. The tool on this page is free, has no page limit, and nothing is uploaded: after the first run, turn Wi-Fi off and it still works.
Make it searchable now
Try it here, nothing is uploaded
- Drop the scanned PDF on the OCR tool above.
- Pick the language of the document under Language. The panel starts from your browser’s language, so change it when the paper is in another one: the pack decides which letters the engine can recognise at all. For a page that mixes 2 languages, set the second under And; it reads more slowly.
- Run it. Each page is drawn at 300 dpi and read on your device by the Tesseract engine, running as WebAssembly in your browser. Nothing leaves your machine.
- Read the report. It gives the average confidence; any page below 60 percent is flagged with a warning, and when 2 or more words were read with low confidence, the report names them and their pages, so you know where to proofread.
- Download. The result is your scan plus the new text layer. Nothing else in the file has changed.
The first run fetches the PDF engine (4.6 MB), the reader (about 1.5 MB) and the language pack you picked (0.7 to 3 MB). All of them come from this site, no third party, and your browser keeps them: after that the tool works offline.
Two behaviours are worth knowing before a long document. Pages that already contain text are left alone, so running OCR over a mixed file does not duplicate what is already searchable; set Pages with text to “read again too” if you want those pages read anyway. And the reading follows the page as your viewer shows it: a page turned with Rotate PDF is read upright. A page whose scan lies sideways and is shown sideways reads badly, so turn it first.
One more thing, since it is the first question people ask: yes, this is actually free, with no cap. The big PDF sites ration it, because their OCR runs on their servers and servers cost money per file: iLovePDF’s free plan reads 1 file of up to 15 MB per task, and Sejda’s free desktop app reads files of up to 10 pages (both checked on 11 October 2026). Ours runs on your device, so there is no per-file cost to pass on, no account, and no watermark.
What OCR adds: an invisible text layer
A PDF page can carry an image and text at the same time, each drawn in its own layer. A scanner only ever produces the image. OCR adds the text: for every word it recognises, it writes that word into the file in an invisible render mode, positioned exactly where the printed word sits on the page.
Because the image draws on top and the text is invisible, the page looks exactly as it did before. But every PDF reader, Acrobat, Preview, Chrome, Edge, your phone, searches and selects the text layer, not the pixels. So Ctrl+F finds “TOTAL”, dragging across the printed line selects the words behind it, and copy puts real text on your clipboard.
The scan itself is not re-encoded, resampled or touched in any way. Our test suite compares the image bytes before and after OCR and requires them to be identical, byte for byte. The file grows a little, because the recognised text has to live somewhere, but text is small: a few kilobytes per page against the hundreds of kilobytes a scanned image takes.
Check the result yourself
Do not take the claims on trust; both of the big ones are checkable in under a minute.
That the file is searchable. Open the result in any reader, pick a word you can see on the page, and press Ctrl+F (Cmd+F on a Mac). The hit lands on the printed word. Then drag across a line: the selection follows the print, because the invisible word underneath has the same position and width as the printed one.
That nothing is uploaded. Run the tool once, then turn Wi-Fi off, or switch the phone to airplane mode, and run it again on another scan in the same language. It works, because the engine and your language pack are already on your device and the reading happens there. A tool that uploaded your scan would stop dead. No upload-and-wait OCR site can offer you that test, which is exactly why we keep suggesting it.
If the result disappoints, the fix is almost never a different button on this page; it is a better scan or the right language pack, covered below.
Read the confidence report
After the run, the report shows a confidence score. It is the engine grading its own guesses, averaged over the pages it read: 95 means it was nearly certain about nearly every word, 70 means it guessed often. The number measures the engine’s certainty, not the truth; it has never seen the original paper, only your scan of it. That makes it a map of where to look, and a good one.
Use it like this:
- 90 and above: trust it. Search and copy freely; spot-check numbers if money or dates depend on them.
- 60 to 90: usable, proofread what matters. Copy the text, but read any figure, name or reference number against the picture before it goes into an email.
- Below 60: the page gets a warning, and the warning means it. At that level, reading the page yourself beats trusting the layer. Usually the cause is a poor scan or the wrong language, both fixable.
Single words get the same treatment. When 2 or more words in a file score under 60 percent, the report counts them, names their pages and quotes a few (on a test page: “Jona,”, “Us”, “pald”), even on a page whose average looks fine. Check names, numbers and dates there first.
One caution: a decent score with the wrong language pack produces tidy nonsense, words that are spelled plausibly and wrong. If the output reads strangely, check the language before blaming the scan.
The 11 languages, and why the pack matters
The engine matches letter shapes against 1 alphabet at a time. Accented letters only exist in their own language packs: run a German letter through the English pack and every umlaut degrades, run Russian through any Latin pack and you get pure noise. Pick the language of the document, and these are the 11 that ship today:
| Language | What its pack brings |
|---|---|
| English | the plain Latin alphabet |
| German | ä, ö, ü and ß |
| French | é, è, ê, ç, œ and friends |
| Spanish | ñ, accented vowels, ¿ and ¡ |
| Italian | à, è, ì, ò, ù |
| Portuguese | ã, õ, ç and the accented vowels |
| Dutch | ë, é and the ij pair |
| Polish | ł, ą, ę, ś, ż, ź |
| Czech | č, ř, š, ž, ě, ů |
| Croatian | č, ć, đ, š, ž |
| Russian | the Cyrillic alphabet, a thing apart |
Each pack is a one-time download of 0.7 to 3 MB from this site, kept on your device afterwards. For a document that mixes 2 languages, pick both: the first under Language, the second under And. Reading takes longer, and the special characters of both survive.
4 things that decide accuracy
OCR quality is set before the engine ever runs, by the scan. Four factors do most of the work:
- Resolution. 300 dpi is the sweet spot, and it is why the tool draws every page at 300 dpi before reading it. A scan made at 150 dpi gives each letter a quarter of the pixels, and small type starts to crumble. Scanning higher than 300 buys almost nothing.
- Flat pages. A curled page or a book spine bends the lines, and a bent line is one the engine splits or merges. Flatten the page under glass or a weight; it is worth more than any software correction after the fact.
- Even light. For phone photos, daylight from the side beats the flash, which drops a glare oval exactly where the text is. Grey toner and faded thermal paper cost accuracy the same way: less contrast, more guessing.
- The right language pack. Free, instant, and routinely worth more than the other 3 combined. The wrong pack fails silently, which is what makes it dangerous.
If a result is mediocre, rescan at 300 dpi with the page flat before trying anything cleverer. A better original beats every trick, and for genuinely poor sources, old faxes, 1-bit archive scans, 6-point receipt type, expect low scores and proofread accordingly.
After OCR: what the text layer unlocks
The searchable file is usually not the end of the job. Everything below works on the OCR result, and everything below runs on your device too.
Search and copy anywhere. The text layer is standard PDF text, so it works in every reader, in Windows Search and Spotlight once the file is indexed, and in any tool that reads PDFs. No getPDF anything required after the download.
Pull the plain text out. When you want the words and nothing else, a paragraph for an email, the text for a records system, extract it:
PDF to textAll the text, in reading order, as UTF-8. Free, runs on your device.Make it a Word document. PDF to Word turns the recognised text into an editable .docx with flowing text, headings and lists. The layout is approximate and the tool says so before you click; the full picture of what survives the trip is in the conversion pillar.
Edit the recognised text in place. The editor can work with the text layer directly, for fixing a recognised date or adding a note on the scan; the editing pillar covers what editing a PDF can and cannot do.
Shrink it if it is huge. OCR adds kilobytes, but the scan itself may be tens of megabytes. That is a compression job, not an OCR job: Compress PDF handles it, and Make a PDF smaller explains the trade-offs.
The honest part
OCR has hard limits, and this page would rather you hit them here than after an afternoon of scanning.
Handwriting mostly does not work. The engine is trained on printed type. Joined-up handwriting comes out with wrong words; the report names the words it was unsure of, but a word it misreads with confidence slips through. Neat block capitals partly survive, digits least of all. For handwritten pages, retyping is usually the honest answer; Handwriting and OCR has our measurements and the alternatives.
A signed scan loses its signature. If the PDF is digitally signed, the searchable copy is a new file, and the tool warns that its signature no longer checks out. Keep the signed original for anything that counts.
Very small print and poor copies lose accuracy. Receipt type at 6 points, a fax relayed twice in 2003, a photocopy of a photocopy: each gives the engine fewer honest pixels per letter, and accuracy falls with them. The confidence report is there to show you where.
The look of the page never improves. OCR adds text behind the picture; it does not touch the picture. A crooked scan stays crooked, a dark scan stays dark, coffee stains survive. If the page needs to look better, rescan it; this tool will not pretend otherwise.
The pixels remain the document. The text layer is the engine’s best reading, not a certified transcript. For anything where a digit matters, a contract figure, an IBAN, a date on a deadline, read the picture. The layer makes the scan findable; your eyes make it trustworthy.
Within those limits, the deal is simple: drop a scan, wait a moment, and leave with a file that looks exactly the same and finally answers Ctrl+F. On your device, free, no account, and provable with the Wi-Fi switch.
Questions
Does OCR change how my scan looks?
No. The recognised text goes into the file as an invisible layer behind the page image, and the image bytes themselves are untouched: our tests compare them byte for byte before and after. Printing, zooming and the look of every page stay exactly the same.
Is my file uploaded anywhere?
No. The OCR engine runs as WebAssembly in your browser, so the reading happens on your device. After the first run has fetched the engine and your language pack from this site, you can turn Wi-Fi off and the tool still works.
Which languages does the OCR support?
11 today: English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Czech, Croatian and Russian. Pick the language of the document, not of your computer, and add a second one for a page that mixes 2. Each pack is a one-time download of 0.7 to 3 MB.
Can it read handwriting?
Mostly not. The engine is trained on printed type, so joined-up handwriting comes out with wrong words. The tool names the words it was unsure of, but a confidently misread word slips through. Neat block capitals partly work. For handwritten pages, retyping is usually the honest answer.
What does the per-page confidence score mean?
It is the engine's own estimate of how sure it was about its guesses, averaged over the page. A page below 60 percent gets a warning, and so do single words under 60 percent, by name: read those against the original before trusting the text layer. A high score on the wrong language pack can still be tidy nonsense, so skim one line.
The tools for this job
Every guide in Scans and OCR
- Handwriting and OCR: what works, what fails, and what to use insteadHandwriting OCR mostly fails and this guide says so up front: what print OCR can read, where block capitals get to, and the honest alternatives that work.
- Make a PDF smaller, with the method that fits your fileMake a PDF smaller with the lever that fits your file: re-encode images, lower DPI, go grayscale or drop pages.