Blog

Scanned PDF vs Digital PDF: Why Text Extraction Fails and How to Fix It

If you've ever tried to copy text from a PDF and couldn't select anything โ€” or pasted a garbled mess of symbols โ€” you've encountered the fundamental divide between digital PDFs and scanned PDFs. Understanding the difference is the key to choosing the right extraction method every time.

The Two Types of PDF

Digital PDF (Text-Based)

Created directly from software โ€” Word, Google Docs, InDesign, LaTeX, or any "Export to PDF" function. The file contains actual character data and text layers alongside the visual rendering. You can select, search, and copy text natively. This is how most modern PDFs work.

Scanned PDF (Image-Based)

Created by scanning a physical document โ€” a printer/scanner that photographs each page and saves the photo as a PDF. The file contains only images of pages. There is no text layer. You cannot select, search, or copy text. The words that appear on screen are just pixels in an image.

A PDF can also be a hybrid: a scanned page with an OCR text layer added on top โ€” this is how Adobe Acrobat's "Make Searchable" feature works, and what most modern multifunction printers do automatically when set to "searchable PDF" mode.

How to Tell Which Type You Have

The fastest test: open the PDF and try to double-click a word. If the word highlights and you can select individual letters, it's a digital PDF. If nothing happens โ€” or clicking selects the entire page like an image โ€” it's a scanned PDF.

A second test: try Ctrl+F (Find). If the search finds your term, the PDF has a text layer. If the search returns zero results for a word that's clearly visible on the page, the text is embedded in an image with no searchable layer.

How OCR Works

OCR (Optical Character Recognition) is the technology that converts images of text into actual machine-readable characters. Modern OCR systems use convolutional neural networks trained on millions of text images to recognise character shapes at the pixel level.

The process, simplified:

  1. Pre-processing: The image is deskewed (rotated to be straight), binarised (converted to pure black and white), and noise-reduced to clean up scan artefacts.
  2. Layout analysis: The engine identifies text regions, separating them from images, tables, and headers/footers. Multi-column layouts are analysed for reading order.
  3. Character recognition: Individual glyphs are identified and matched against character models. Context (surrounding characters, language model probabilities) helps resolve ambiguous shapes.
  4. Post-processing: The recognised text is run through a language model to catch obvious errors (e.g., "tlie" being corrected to "the").

OCR Accuracy: What to Realistically Expect

99%+

Clean scan, standard font, 300+ DPI, good contrast

95โ€“98%

Slightly skewed, minor noise, 200 DPI, printed text

85โ€“94%

Low DPI, faded ink, unusual fonts, background patterns

<85%

Handwriting, heavily damaged, very low resolution, decorative fonts

At 99% accuracy on a 500-word page, you'd expect roughly 5 errors โ€” acceptable for most purposes. At 95%, that's 25 errors per page. For legal or medical documents, even 5 errors per page requires careful proofreading.

Tips for Better OCR Accuracy

FactorRecommendationWhy It Matters
Scan resolution300 DPI minimum; 400+ for small textLower DPI means character shapes are pixelated and harder to distinguish
ContrastBlack text on white backgroundOCR engines struggle with light ink, coloured paper, or watermarks behind text
SkewScan straight โ€” correct skew if crookedRotated text dramatically reduces accuracy; most tools auto-deskew
Language settingMatch the document's languageLanguage models assist recognition; wrong language = wrong corrections
Font typeStandard serif/sans-serif print is bestDecorative, handwritten, or very compressed fonts degrade accuracy significantly
DamageRemove fold lines and shadows if possibleScan artefacts appear as spurious characters or break character shapes

Common OCR Errors to Watch For

OCR OutputLikely Actual CharacterWhy It Happens
rnmThe letters r and n together look identical to m at lower resolutions
0OZero and capital O are visually similar, especially in older fonts
1l or IDigit 1, lowercase L, and capital I are virtually identical in many fonts
5SEspecially common in faded print or stylised fonts
cldThe letters c and l together resemble d in certain fonts at low DPI
iiuDotted letters together can be misread as u, especially in cursive-like fonts
Google Drive free OCR: Upload a scanned PDF to Google Drive, right-click it, and choose "Open with Google Docs." Google automatically runs OCR and converts the PDF into an editable Doc. For many standard scanned documents, this produces excellent results at no cost โ€” and no file size limit concerns.

Frequently Asked Questions

Can OCR read handwriting?

Modern OCR tools with handwriting recognition models (like Google Vision AI or Microsoft Azure OCR) can read clear, neat handwriting with reasonable accuracy โ€” often 85โ€“95% on good quality scans. Cursive and irregular handwriting remains significantly harder, typically below 85% accuracy. For critical handwritten documents, manual transcription is still more reliable.

Why does my OCR output have correct words but in the wrong order?

Multi-column PDFs confuse OCR layout analysis. The engine may read column 1 line 1, then jump to column 2 line 1, then back to column 1 line 2 โ€” producing text that reads correctly but in the wrong paragraph order. Use an OCR tool with "column detection" or manually specify the reading zone. Re-ordering the output is sometimes necessary for complex layouts.

Is a digitally created PDF always better than a scanned one?

For text extraction, yes โ€” digital PDFs allow perfect extraction with no errors. However, some older or legacy documents only exist as physical papers, and some content (signatures, stamps, handwritten annotations) can only be captured by scanning. For archival purposes, a high-quality scan with an OCR text layer is the best of both worlds.

Convert Your PDF to Text โ€” Free

Supports both digital and scanned PDFs. Upload and download your extracted text in seconds.

Open PDF to Text Converter โ†’