<link rel="stylesheet" href="/assets/fonts/jetbrains-mono/jetbrains-mono.css" />
All posts

OCR: how to extract text from scans and photos (and why it sometimes fails)

You receive a PDF, open it, try searching for a word: no results. You try selecting a line with the mouse: nothing selects. The document is right there, perfectly readable to your eyes, but to the computer it's a photograph. This is where OCR — optical character recognition — comes in, and the moment most people discover that the result depends far more on the source image than on the software.

What OCR actually does (and why it doesn't "read")

An OCR engine doesn't understand text: it reconstructs it. The process goes through mechanical stages — locating text areas, straightening the image, isolating individual lines, then words, then characters — and only at the end does it compare each shape against the letter models it knows, picking the most probable one.

This explains the typical errors, which look absurd at first glance: rn read as m, 0 confused with O, 1 with l. These aren't comprehension errors, they're visual similarity errors. A human corrects them automatically thanks to context; an OCR engine, which reasons about shape before meaning, does not.

Image quality matters more than the software

Here's the counterintuitive part: switching OCR engines improves little, improving the scan improves a great deal. The factors that genuinely weigh, in order of impact:

FactorProblemRemedy
ResolutionBelow 200 DPI small characters smear togetherScan at 300 DPI, the de facto standard
ContrastYellowed paper or a dim photo: letters and background blur into each otherBlack and white, or raise contrast before processing
SkewCrooked lines: the engine segments them wronglyStraighten the image, even by a few degrees
Shadows and curvaturePhoto of an open book: text warps towards the spineShoot with diffuse light, page flattened

A photo snapped quickly on a phone, crooked and with your hand's shadow across it, produces mediocre results with any engine. It's worth redoing the capture: the camera document scanner applies straightening, cropping and contrast filters before the file even reaches recognition, and that's almost always the difference between usable text and text you'll retype by hand.

The confidence score: the most useful figure, and the most ignored

Every processed page gets a confidence value — how reliable the engine considers its own reading. It's the most valuable piece of information in the whole process, and almost nobody looks at it.

The practical rule: above 90% the text is generally reliable and just needs proofreading; between 70% and 90% there are scattered errors to fix; below 70% you're better off redoing the capture than wasting time correcting. A low score on one page out of twenty almost always points to a physical problem with that page — a fold, a shadow, a stain — not a limitation of the system.

How to extract the text, step by step

  1. Open the free online OCR — no sign-up.
  2. Upload the files: PNG, JPG, WEBP, BMP, TIFF images or PDFs, including scanned ones. Up to 15 MB per file and 30 files at a time.
  3. Select the right language before starting. This is the step with the biggest impact on accuracy: English, Italian, Spanish, French, German, Portuguese and Albanian are available.
  4. Run the recognition and check the confidence score page by page.
  5. Copy the text, or send it to the workspace to pass it to another tool — typically a summary or a translation — without re-uploading anything.

Where OCR fails systematically

  • The wrong language ruins everything. An Italian text processed as English loses accents and apostrophes systematically, because the engine looks for shapes that don't exist in Italian.
  • Tables become linear text. Cells are read one after another: the values are there, the row-column relationship isn't.
  • Multiple columns get mixed up. In newspapers and magazines the text may be reconstructed across columns rather than down them.
  • Italics and decorative fonts have far higher error rates than a regular typeface.
  • Handwriting is a different technology. Classic OCR recognises printed characters; handwriting requires different models and remains unreliable.

Frequently asked questions

Which formats can I upload?

Images in PNG, JPG, WEBP, BMP and TIFF, plus PDFs, including scanned ones containing no selectable text. The limit is 15 MB per file, up to 30 files per run.

Which languages does it recognise?

Seven: English, Italian, Spanish, French, German, Portuguese and Albanian. The language must be chosen before processing, because it guides the engine in recognising accented characters and combinations typical of that language.

How do I know whether a PDF needs OCR?

Try selecting a word with the mouse or searching for a term with Ctrl+F. If selection doesn't work and the search finds nothing even for words visibly on the page, the PDF contains images and needs OCR.

Can I use OCR on a photo taken with my phone?

Yes, and it's one of the most common uses. The result improves noticeably if the page is framed straight, well lit and free of shadows: going through the scanner first, which straightens and boosts contrast, almost always yields cleaner text.


In summary

OCR turns an image of text into real text, and it's the prerequisite for almost everything else: searching inside a document, translating it, summarising it, converting it. Accuracy depends above all on three things, in this order: capture quality, correctly selected language, and checking the confidence score before trusting the result. The free online OCR accepts images and PDFs up to 15 MB in seven languages; if the document is still on paper, start from the scanner, and from there the text can move on to the other PDF and AI tools.

💬 Reader notes

0 notes

Write a note

Share your opinion, a suggestion or a compliment

Latest notes

No notes yet. Be the first to comment!