The visible page is only one layer of the story
A PDF viewer ultimately paints a page. That paint may come from text drawing operators, vector paths, raster images, form objects, annotations, or a mixture of all of them. A document can therefore look like a page of text while containing no actual text objects at all. That is common with scans: the page is simply a photograph of paper.
The distinction matters because tools operate on what the PDF contains, not only on what your eyes see. Search needs characters or a searchable OCR layer. Copy and paste needs character mappings. Structural compression can preserve text and vectors without touching their appearance, while a scan may need image recompression to shrink substantially.
Three common kinds of PDF page
| Page type | What is stored | What usually works |
|---|---|---|
| Born-digital text PDF | Text objects, fonts, vectors and images | Search, selection, copy, structural compression and often higher-quality zoom |
| Image-only scan | One or more raster images | Viewing and image processing; text search requires OCR |
| Searchable scan | Page image plus an OCR text layer | Looks like the scan, but search/copy may work because invisible or transparent text is positioned over it |
Quick tests you can do without specialist software
First try dragging across a sentence. If individual words highlight cleanly, the page probably contains text. If the whole page behaves like one picture, it is probably image-only. Then use the viewer's search function for a distinctive word. A searchable scan may pass the search test even though its visible page is still an image.
Copying into a plain-text editor gives another clue. Clean, correctly ordered text suggests a usable text layer. Garbled characters, missing spaces or bizarre reading order can indicate a bad character map, unusual font encoding or low-quality OCR. That is why “I can highlight it” is not the same as “the text structure is good.”
Why OCR does not turn a scan back into the original document
OCR estimates characters from pixels. It does not recover the original word processor file, original font metrics, paragraph styles or semantic structure. A good OCR layer can make a scan searchable while leaving the photographed page untouched, but recognition errors remain possible. Small text, skew, compression artefacts, handwriting, tables and unusual typefaces all make recognition harder.
Why this changes compression choices
If a file is mostly text and vectors, converting every page to JPEG can make the file larger while also destroying selectable text. Structural optimisation is usually the safer first attempt because it can recompress streams and reorganise objects without redrawing the document as pictures. For a scan where almost all bytes are high-resolution images, image downsampling may have a much larger effect.
Why this changes editing choices
Adding an annotation over a text PDF does not necessarily require flattening the original page. A careful editor can preserve the source page and overlay the new content. Destructive changes are different. Proper redaction, whiteout used as replacement, and direct replacement of existing text may require a flattened output path so the material underneath is not recoverable.
A practical workflow
- Try text selection and search.
- If the page is image-only, OCR a copy if searchability matters.
- Keep an untouched original when the document matters.
- Use structural compression first for text-heavy files.
- Use image-based compression only when you accept rasterisation or the source is already image-based.
- After conversion, verify search, copy, page count and visual appearance instead of assuming success from the download alone.