How to make a scanned PDF searchable with OCR

You scanned a document, or someone sent you one, and now you cannot search it, copy a number out of it or even select a word. That is because a scan is a photograph of text, not text. OCR, optical character recognition, reads the picture and gives the words back to the file.

What OCR does to a PDF

The page stays exactly as it was: the same picture, in the same place. OCR reads the words in it and adds them as an invisible layer on top, each word where it sits in the picture. You see the original scan, but when you search, select or copy, you are using the layer underneath.

Make a scan searchable

  1. Add the scanned PDF to OCR PDF.
  2. Tick the language of the text. For a page in two languages, tick both, up to three in all.
  3. Choose Standard quality. Use Detailed for small print, or Fast for a long document that is clearly printed.
  4. Press Read the text. A page takes a few seconds, and a progress line shows how far it is.
  5. Download the result, open it, and press Ctrl+F to search for a word you can see on the page.

How to get an accurate result

  • Start from the best scan you can make: straight, evenly lit and at 200 to 300 dpi. Phone photos work much better after Scan to PDF has straightened them.
  • Make sure the pages are the right way up. Turn them with Rotate PDF first if they are not.
  • Tick only the languages that are really on the page. Extra languages slow the reading and make it less sure.
  • Do not expect handwriting to work. OCR is for printed text.
  • Read the summary: it names the pages the engine was less sure of, and those are worth a look.

What you can do afterwards

  • Search a long document for one name or number.
  • Copy a figure or a paragraph without retyping it.
  • Black out private details with Redact PDF, whose search can now find them in the scan.
  • Convert the document to Word, Excel or plain text with PDF to Word, PDF to Excel and PDF to Text.

Try the tools