Optical character recognition (OCR) turns images of text — scans, photos, screenshots — into machine-readable text.
How Modern OCR Works
Deep-learning OCR typically has two stages: text detection finds regions containing text, and text recognition reads the characters in each region. Some modern models, including multimodal language models, do both and understand layout at the same time.
Beyond Plain Text
Many business documents need more than raw text:
- Layout analysis: columns, headings, reading order.
- Tables: preserving rows and columns.
- Key–value extraction: invoice number, date, total.
- Handwriting recognition, which remains harder than print.
Where OCR Fails
Low resolution, blur, skewed photos, unusual fonts, stamps and signatures overlapping text, faded print, complex tables and poor contrast.
Improving Results
- Scan at adequate resolution (around 300 dpi for documents).
- Correct rotation and perspective; improve contrast.
- Specify the language(s).
- Use a model trained for your document type where possible.
Validate Extracted Data
Check extracted fields against rules — dates that exist, totals that add up — and route low-confidence results to a person.
Measure Accuracy
Use character error rate or word error rate on a sample of your real documents, and field-level accuracy for key–value extraction.
Privacy
Scanned documents often contain personal data. Choose tools whose data handling meets your obligations.