Article · Scanning

Scanned PDFs vs Real PDFs: Why Your 'PDF' Is Just Pictures

You try to search the contract for "termination clause" and the reader finds nothing. You try to copy an address and get a gray box. Your "PDF" looks like a document but behaves like a photograph — because that is exactly what it is.

Two kinds of PDF, and how to tell them apart

A real PDF stores text as text. Every letter is a character the computer understands, which is why you can select sentences, search for words, and copy an address into an email. These files are created when software — Word, Google Docs, an invoicing app, a web browser's Print to PDF — writes the document directly.

A scanned PDF stores each page as an image. A scanner or phone camera photographed the paper, and the PDF is a wrapper around those photographs. To the computer, the page contains no words at all — just pixels that happen to look like words to your eyes.

The one-second test: try to select a sentence with your mouse. If you can highlight and copy the words, you have a real PDF. If the selection grabs the whole page as one block — or nothing at all — you are looking at pictures. This single test explains nearly every mysterious PDF behavior people complain about.

Why the difference matters: search, copy, and file size

The gap between the two kinds shows up in three places that affect your daily life:

Accessibility matters too: screen readers used by visually impaired colleagues can read real text aloud, but a scanned page is silence. If your document needs to reach everyone, it needs real text.

What OCR does — and what it cannot do

OCR — optical character recognition — is the bridge between the two kinds. It examines the images in a scanned PDF, recognizes the letters in them, and adds an invisible layer of real text behind each page. The pages still look exactly like the scans, but suddenly Ctrl+F works and sentences can be copied.

On clean, printed text, modern OCR is remarkably accurate — well over 99% of characters correct. It struggles where humans struggle: handwriting, faded print, skewed pages, coffee stains, and unusual fonts. It also guesses, and its guesses can be wrong in ways that look right at a glance — "rn" misread as "m" is a classic. For anything legally or financially important, proofread the OCR result rather than trusting it blindly.

And one hard limit: OCR cannot recover what was never there. It reads the scan and adds a best-effort text layer — it cannot restore the original digital document with its formatting, fonts, and structure. An OCR'd scan is a searchable picture, not a reborn original.

The best OCR is the scan you never needed. Before running OCR on a pile of scans, ask whether the original digital file still exists — in your email, in the app that generated it, from the sender. A genuine text PDF beats even perfect OCR every time.

Practical fixes for the PDFs you already have

Different problems, different fixes:

  1. File too big to send? Scanned PDFs are the number-one cause of "attachment too large." Run the file through the Compress PDF tool — scans shrink dramatically because their page images have enormous room to slim down. Our guide on reducing PDF file size walks through the quality settings.
  2. Need to search or copy text? Run OCR on the file. Most scanner apps and desktop PDF readers include it; the result keeps the familiar scanned look while making every word findable.
  3. Building new scans from paper? Photograph pages well — flat, evenly lit, squared to the frame — and run your scanner app's text-recognition option at capture time. Good input makes everything downstream better. See our article on turning phone photos into a PDF for the technique.
  4. Combining scans with real documents? Merging works fine across both kinds — but remember the merged file inherits the scans' weight and unsearchability. Compress after merging if size is an issue.

Scan FAQs

How can I tell if my PDF is a scanned image or real text?

Try to select a sentence with your mouse. If you can highlight and copy the words, it is a real text PDF. If the selection grabs the whole page as one block — or nothing at all — the page is an image.

Why are scanned PDFs so much larger than normal PDFs?

A text PDF stores characters — a few bytes per word. A scanned PDF stores a full-page photograph at high resolution, often a megabyte or more per page. Fifty pages of scans can easily weigh 100 MB where the same text as characters would be under 1 MB.

What does OCR do to a scanned PDF?

OCR — optical character recognition — reads the text in the page images and adds an invisible text layer underneath, so you can search and copy words while the pages still look like the original scans. Good OCR is very accurate on clean print; handwriting and blurry scans are harder.

Can I convert a scanned PDF back into a real text PDF?

You cannot recover the original digital text — it was never in the file. But running OCR adds a searchable text layer, and re-exporting the document from its original source (Word, email, accounting software) gives you a genuine text PDF whenever the source still exists.

Scanned PDF too heavy to send?

Shrink bloated scans for email — free, private, on your device.

Open the Compress PDF tool