Search scanned pages
Recognised words are added to the PDF as an invisible text layer while the original scan remains visible.
Turn scanned PDF pages into searchable documents. NoblePDF recognises the words on each page in your browser and adds a searchable text layer without replacing the original page image.
Make image-based pages searchable and selectable.
or drop a PDF here
OCR converts text that exists only as pixels into a machine-readable text layer so you can search, copy and index scanned documents.
Recognised words are added to the PDF as an invisible text layer while the original scan remains visible.
After OCR, compatible PDF readers can search and select recognised text that was previously just an image.
PDF rendering, recognition and PDF rebuilding happen in your browser. The OCR runtime may download code and language data.
Select a scanned or image-based PDF from your device.
Choose the document language and let NoblePDF recognise each required page.
Save a new searchable copy while keeping the original page appearance.
NoblePDF’s OCR workflow starts by rendering the PDF page with PDF.js and passing the resulting image data to Tesseract in the browser. The recognised words can then be represented as a text layer while the original visual page remains in place. This makes a scan searchable without requiring the page to be replaced by newly typeset text.
The current runtime bundles English, Spanish, French, German, Italian and Portuguese language data. Choosing the language that best matches the page can improve recognition, especially for accented characters and common word patterns. Mixed-language pages, handwriting, equations and stylised fonts remain more difficult.
Scanned-pages mode is useful when a PDF contains a mixture of real text pages and images. Re-running OCR over a page that already has a good text layer can create duplicate or misaligned text, so the tool can leave those pages alone. Use all-pages mode when there is a specific reason to replace or augment existing text.
OCR accuracy is not a legal or financial guarantee. Check names, dates, amounts and identifiers against the page image before relying on them. A clean-looking output can still contain substitutions such as 0/O, 1/l or punctuation mistakes.
Because recognition is CPU-intensive, long high-resolution scans can take significant time and memory. Splitting a very large scan into sections can improve reliability on mobile devices. The supported OCR path keeps the selected document and recognition work in the browser instead of uploading page images to a NoblePDF OCR service.
For the supported OCR workflow, your PDF stays in the browser. PDF.js renders pages locally and Tesseract.js performs recognition in the browser. NoblePDF serves the OCR engine and supported language data from its own origin.
The original PDF page content remains in place. NoblePDF adds an invisible text layer over recognised words rather than rasterising the document again.
Recognition depends on scan resolution, blur, skew, handwriting, unusual fonts and language selection. Always review important extracted text.
The default Scanned pages only mode skips pages that already contain embedded text. Choose All pages only when you specifically want OCR applied everywhere.
Searchable OCR text is different from tagged accessibility structure and different again from native born-digital text. These pages explain the boundary.