A scanned PDF is a picture of a page, so its words can't be selected, searched or copied until OCR (optical character recognition) reads them. To extract text from a scanned PDF, upload the file to a text recognition tool, let it recognise the characters, then copy or download the text. It takes a few minutes, which is a lot less than retyping.
It usually happens at a bad moment. You need one clause from a signed contract for an email, the reference numbers from a supplier invoice for your accounts, or a paragraph from an old report that only exists on paper. The PDF opens, everything is legible, and the cursor slides straight over the text. This guide explains why, how to check what kind of file you have, and how to get the words out cleanly.
Why can't you select text in a scanned PDF?
When a scanner or a phone captures a page, it records what the page looks like, not what it says. The result is saved as a PDF, but each page inside the file is essentially one image. Your screen shows letters; the file only holds pixels.
That is why the document behaves normally until you need its content. It opens, it prints, it can be shared. Then a search for a name finds nothing, and selecting a sentence either grabs the whole page or nothing at all.
OCR is what bridges the gap. It analyses the shapes in the image, identifies them as letters, numbers and punctuation, and turns them into real text you can work with. If you want to know more about the technology itself, our explainer on optical character recognition covers how it works.
How do you know if your PDF contains text or just an image?
Two checks take less than a minute and stop you from running recognition on a file that doesn't need it.
First, search for a word you can clearly see on the page. If your PDF reader finds it, the text is already there and you can copy it directly. If the search returns nothing, you are looking at an image.
Second, zoom in to 400% or more. Real text stays sharp at any zoom level because it is drawn from font data. Scanned text turns blurry or blocky, because you are enlarging a photo.
Some files mix both. A report exported from a word processor can include a scanned signature page, for example. In that case, only the scanned pages need recognition.
How to extract text from a scanned PDF, step by step
Online text recognition runs in your browser, so there is nothing to install. With PDFSmart, the process looks like this:
- Keep a copy of the original scan, so you can run it again if the first result isn't clean.
- Open the
- Let the tool recognise the text, then select the passages you need.
- Copy the text into your document, or download the extracted text to keep it as a separate file.
- Proofread names, figures and dates before you use them anywhere that matters.
That last step deserves a minute. OCR is reliable on clean printed pages, but a smudged digit in an invoice total or a bank account number is exactly the kind of error that slips through unnoticed. Reading the extracted figures against the scan is quicker than sending a correction later.
Which method should you use?
Recognition is not the only way to get text out of a scan. Here is how the usual options compare.
Retyping by hand is fine for a single line, such as a reference number. Beyond a few sentences, it becomes slow and errors creep in.
Copying text from a screenshot with a feature built into your operating system works for one short passage on one page. It isn't available on every system, and you can only capture one screen at a time.
Running OCR on the whole file with an online tool is the practical option for full pages of text you need to reuse. It needs a readable scan, and handwriting is recognised less reliably than print.
Rescanning the page is the answer when the scan itself is blurry, tilted or shadowed. It only works if you still have the paper original.
How to get a cleaner result from OCR
Most recognition errors come from the scan, not from the recognition. If your text comes back with odd characters or words stuck together, look at the source file before anything else.
Scan at 300 DPI or more: below that, small characters lose the detail recognition relies on. Keep the page flat and straight, since tilted lines are harder to read. Light it evenly, without the shadow of your hand or phone across the paper. And expect printed text to come back far more accurately than handwriting, which remains difficult for OCR in general.
If the original is still on your desk, rescanning often beats correcting. From your phone, the PDFSmart mobile scanner turns a photo of the page into a PDF that you can then run through text recognition.
What can you do with the text once it's extracted?
Once the words are out of the image, they behave like any other text. You can paste a clause into an email, drop figures into a spreadsheet, or rebuild the document in your word processor and update what has changed since the original was printed.
If that rebuilt document has to go back out as a PDF, convert your Word file to PDF once it is final. You end up with a lighter file whose text can be searched and copied, which the scan could never offer.
A habit worth keeping: when you know you will need a document's content later, extract the text right after scanning, while the paper is still within reach.
In short
A scanned PDF holds images, not text, which is why nothing can be selected or searched. Check the file with a quick search, run the scan through text recognition, then proofread the figures before reusing them. A clean scan at 300 DPI does most of the work for you.
FAQ
What does OCR stand for?
OCR stands for optical character recognition. It reads the shapes in an image, identifies them as letters, numbers and symbols, and converts them into text you can copy, search and edit.
Why is my extracted text full of errors?
The quality of the scan is almost always the cause. Low resolution, a tilted page, shadows or faint print make characters harder to recognise. Rescan at 300 DPI or more with even lighting, then run the recognition again.
Can OCR read handwriting?
Recognition works best on printed text. Handwriting varies too much from one person to another to be read with the same reliability, so expect more corrections and always check the result against the original.
Can I extract text from a photo of a document?
Yes. Besides PDF, the PDFSmart text recognition tool accepts common image formats such as JPG, PNG, TIFF and WEBP, so a photo of a document is processed the same way as a scan.
Do I need to install software to extract text from a scanned PDF?
No. Online text recognition runs in your browser: you upload the file, let the tool recognise the text, then copy it or download it.