If you've ever tried to copy text from a PDF and couldn't select anything โ or pasted a garbled mess of symbols โ you've encountered the fundamental divide between digital PDFs and scanned PDFs. Understanding the difference is the key to choosing the right extraction method every time.
Created directly from software โ Word, Google Docs, InDesign, LaTeX, or any "Export to PDF" function. The file contains actual character data and text layers alongside the visual rendering. You can select, search, and copy text natively. This is how most modern PDFs work.
Created by scanning a physical document โ a printer/scanner that photographs each page and saves the photo as a PDF. The file contains only images of pages. There is no text layer. You cannot select, search, or copy text. The words that appear on screen are just pixels in an image.
A PDF can also be a hybrid: a scanned page with an OCR text layer added on top โ this is how Adobe Acrobat's "Make Searchable" feature works, and what most modern multifunction printers do automatically when set to "searchable PDF" mode.
The fastest test: open the PDF and try to double-click a word. If the word highlights and you can select individual letters, it's a digital PDF. If nothing happens โ or clicking selects the entire page like an image โ it's a scanned PDF.
A second test: try Ctrl+F (Find). If the search finds your term, the PDF has a text layer. If the search returns zero results for a word that's clearly visible on the page, the text is embedded in an image with no searchable layer.
OCR (Optical Character Recognition) is the technology that converts images of text into actual machine-readable characters. Modern OCR systems use convolutional neural networks trained on millions of text images to recognise character shapes at the pixel level.
The process, simplified:
Clean scan, standard font, 300+ DPI, good contrast
Slightly skewed, minor noise, 200 DPI, printed text
Low DPI, faded ink, unusual fonts, background patterns
Handwriting, heavily damaged, very low resolution, decorative fonts
At 99% accuracy on a 500-word page, you'd expect roughly 5 errors โ acceptable for most purposes. At 95%, that's 25 errors per page. For legal or medical documents, even 5 errors per page requires careful proofreading.
| Factor | Recommendation | Why It Matters |
|---|---|---|
| Scan resolution | 300 DPI minimum; 400+ for small text | Lower DPI means character shapes are pixelated and harder to distinguish |
| Contrast | Black text on white background | OCR engines struggle with light ink, coloured paper, or watermarks behind text |
| Skew | Scan straight โ correct skew if crooked | Rotated text dramatically reduces accuracy; most tools auto-deskew |
| Language setting | Match the document's language | Language models assist recognition; wrong language = wrong corrections |
| Font type | Standard serif/sans-serif print is best | Decorative, handwritten, or very compressed fonts degrade accuracy significantly |
| Damage | Remove fold lines and shadows if possible | Scan artefacts appear as spurious characters or break character shapes |
| OCR Output | Likely Actual Character | Why It Happens |
|---|---|---|
| rn | m | The letters r and n together look identical to m at lower resolutions |
| 0 | O | Zero and capital O are visually similar, especially in older fonts |
| 1 | l or I | Digit 1, lowercase L, and capital I are virtually identical in many fonts |
| 5 | S | Especially common in faded print or stylised fonts |
| cl | d | The letters c and l together resemble d in certain fonts at low DPI |
| ii | u | Dotted letters together can be misread as u, especially in cursive-like fonts |
Modern OCR tools with handwriting recognition models (like Google Vision AI or Microsoft Azure OCR) can read clear, neat handwriting with reasonable accuracy โ often 85โ95% on good quality scans. Cursive and irregular handwriting remains significantly harder, typically below 85% accuracy. For critical handwritten documents, manual transcription is still more reliable.
Multi-column PDFs confuse OCR layout analysis. The engine may read column 1 line 1, then jump to column 2 line 1, then back to column 1 line 2 โ producing text that reads correctly but in the wrong paragraph order. Use an OCR tool with "column detection" or manually specify the reading zone. Re-ordering the output is sometimes necessary for complex layouts.
For text extraction, yes โ digital PDFs allow perfect extraction with no errors. However, some older or legacy documents only exist as physical papers, and some content (signatures, stamps, handwritten annotations) can only be captured by scanning. For archival purposes, a high-quality scan with an OCR text layer is the best of both worlds.
Supports both digital and scanned PDFs. Upload and download your extracted text in seconds.
Open PDF to Text Converter โ