Skip to content
files.co

How to OCR a Scanned PDF and Make It Searchable

Turn a scanned PDF into a searchable, selectable document. A step-by-step guide to running OCR in your browser, no upload required.

DCDavid Carrero · · 6 min read

You scanned a stack of paper into one PDF, and now you can’t do anything with it. No searching for a word, no copying a sentence into an email, no highlighting a clause in a contract. That’s because a scan isn’t text, it’s a photograph of text. Your PDF reader has no idea what the words say, it just knows there’s an image on the page.

The fix is OCR, optical character recognition. It looks at the shapes on the page, figures out which letters they are, and adds a real, invisible text layer underneath the image. The page still looks exactly the same. But now you can search it, select from it, and copy from it like any normal document.

Why this actually matters

Think about the last time you needed one specific detail out of a scanned file: an invoice number, a clause in a lease, a date on a certificate. Without OCR, you’re stuck scrolling and squinting through every page by eye. With it, Ctrl+F finds the line in half a second.

It matters even more at scale. If you’re digitizing years of paper records, a folder of scans without OCR is basically a filing cabinet you photographed, still hard to search, still hard to reuse. Add a text layer and the same folder becomes something you can actually query, index, and pull data from.

There’s a compliance angle too. Some document standards, like PDF/A for long-term archiving, expect real text rather than flat images wherever possible. A scan that’s been through OCR is simply a better-behaved file, easier to archive, easier to hand off, easier to trust five years from now.

How OCR works, briefly

The tool analyzes the pixels of each scanned page and recognizes patterns that match letters and words, the same basic idea behind the technology that reads license plates or sorts mail. Once it’s confident about a word, it places matching, selectable text exactly where that word sits on the page, positioned behind the image so nothing visually changes.

This is what people mean by a “text layer”: an invisible skeleton of real characters sitting under the picture, there for your PDF viewer’s search and copy functions to find, but never seen on screen.

Accuracy depends on scan quality. A crisp 300 DPI scan of a clean printed page comes out close to perfect. A blurry photo of a handwritten note taken at an angle will produce more errors, OCR reads shapes, not intent, so it does best with clear, well-lit, reasonably straight scans.

Doing it step by step

  1. Open the OCR tool in your browser. No account, no install.
  2. Upload your scanned PDF. If your document has multiple languages, check whether a language option is available and pick the right one, it improves recognition noticeably.
  3. Run the OCR process. The tool reads each page and builds the text layer behind the scanned image.
  4. Download the result. You’ll get back the same PDF, visually unchanged, now searchable and selectable.
  5. Test it. Open the file, hit Ctrl+F (or Cmd+F on a Mac), and search for a word you know is on the page. If it jumps straight to it, you’re done.

That’s the whole process. No settings to fight with, no format conversions, just a file that goes in looking one way and comes out working better.

What makes OCR accurate, and what wrecks it

Factor Good result Poor result
Resolution 300 DPI Below 150 DPI
Source A flatbed scan A phone photo at an angle
Type Printed text Handwriting
Contrast Black on white Faded ink, coloured or patterned background
Straightness Square to the page Skewed, curled at the spine
Language Set to match the document Left on the wrong language

What to do if a few words come out wrong

What to do if a few words come out wrong

OCR is very good, not perfect. Faded ink, unusual fonts, stray coffee stains, they can all trip it up on the odd word here and there. If accuracy matters a lot for a specific document, like a legal contract, it’s worth spot-checking the searchable text against the original page. For most everyday use, a document that’s 99% searchable is a massive upgrade over one that’s 0% searchable.

A note on privacy

Everything above happens directly in your browser. The scanned pages, whatever’s written on them, personal records, financial statements, medical forms, never leave your device to reach a server. There’s nothing to upload, so there’s nothing sitting on someone else’s storage waiting to be a problem later. You open the file, the processing runs locally, you download the result. That’s the whole trip.

If you’re regularly digitizing paper, running everything through OCR once is worth the five minutes. Every future search, every copy-paste, every “where did I put that clause” moment gets faster from here on.

Frequently asked questions

How do I make a scanned PDF searchable?

Run OCR on it. The engine reads the picture of each page, recognises the characters and writes them into the file as an invisible text layer under the image. The page looks identical afterwards — you have not altered the scan — but the document can now be searched, selected and copied, and a screen reader can read it aloud.

Why is my OCR result full of mistakes?

Usually the scan, not the engine. Below about 150 DPI there is not enough detail to tell similar characters apart; a photo taken at an angle, a page curled at the spine, faded ink or a patterned background all cut accuracy hard. Rescanning at 300 DPI, square and flat, fixes more than any setting will. Handwriting is a separate problem and no general OCR handles it well.

Does OCR change how my document looks?

No. The recognised text goes in as an invisible layer underneath the existing image, so the page you see is the same page you scanned. That is deliberate: OCR is not re-typesetting your document, it is annotating it with what it read, which is why a bad recognition never disfigures the page — it just makes the search unreliable.

Is my scan uploaded for OCR?

Not here. The recognition engine runs in your browser through WebAssembly, so the pages are processed on your own machine. Given that the things people OCR are old contracts, medical records and correspondence, the alternative — uploading a stack of scanned personal documents to a server to be read by software — is a lot to accept for a convenience.

Explore by category