OCR

What OCR Can and Cannot Recover From a Scanned PDF

OCR can make a scan searchable and reusable, but it is recognition, not perfect recovery of the original document.

A scan may contain no real text at all

Many “PDF documents” from scanners are simply page-sized images stored inside a PDF container. A viewer can display them, but searching for a name returns nothing because there are no character objects to search. Optical character recognition analyses the rendered pixels and predicts which characters and words they represent.

What NoblePDF adds

NoblePDF renders pages locally with PDF.js and uses Tesseract to recognise supported languages in the browser. The OCR workflow can add an invisible text layer aligned with the page so the original visual scan remains visible while recognised text becomes searchable/selectable. The current package includes six language datasets: English, Spanish, French, German, Italian and Portuguese.

Resolution matters

Tiny characters give an OCR engine very little information. High-resolution scans generally provide clearer letter shapes, but enormous images cost more memory and processing time. Blur, JPEG artefacts, shadows, folded pages and camera perspective can all reduce accuracy even when the nominal pixel count is high.

Language selection matters too

Recognition models use language-specific patterns. An English model confronted with accented French or German text can make substitutions that a better-matched language model would avoid. Mixed-language documents remain harder because one page may contain several vocabularies, abbreviations and proper names.

Layout is not the same as text

OCR can recognise words without perfectly reconstructing tables, columns, footnotes or reading order. A page with two newspaper columns may produce correct words in the wrong sequence. Handwriting, mathematical notation and decorative fonts are additional challenges. If you need an editable Word document, OCR is only one stage of reconstruction.

When to skip pages that already have text

If a page already contains a good embedded text layer, running OCR again can create duplicate or misaligned text. NoblePDF therefore offers a scanned-pages-oriented mode that can leave pages with existing text alone. Use all-pages OCR only when you understand why the existing text is inadequate.

Review important data

Never assume an OCR engine correctly captured names, dates, account numbers, amounts or legal clauses. Compare important values against the page image. OCR is extremely useful for search and accessibility workflows, but the result remains machine recognition and can contain confident-looking errors.

Privacy and performance

Local OCR keeps the page images in the browser for the supported workflow, but it uses significant CPU and memory. Large scanned documents can take time. Splitting a huge PDF and processing logical sections can reduce the chance of a mobile browser terminating the job.

A practical OCR quality checklist

Check OCR quality against the page image rather than trusting the extracted text alone. Names, account numbers, dates, decimal points and low-contrast characters deserve particular attention because a single recognition error can change meaning while still producing plausible-looking text.

When OCR output will be searched or copied, test several representative pages: clean typed text, a dense table, a skewed scan and the poorest page in the file. This gives a better picture of practical accuracy than testing only the first page. If the source language is not supported by the selected OCR model, recognition quality can fall sharply.

Preserve the original scan when accuracy matters. The OCR-derived copy is useful for search and reuse, but the page pixels remain the primary visual evidence. Comparing the result in a second PDF viewer also confirms that the output opens normally outside the tool that produced it.

For archival or evidentiary scans, preserve the untouched image-based PDF alongside the OCR-enhanced copy. The recognised text layer improves discovery, but the original pixels remain the best reference when a character is disputed.