NoblePDF Lab Note

Lab Note: Does PDF Compression Preserve Searchable Text?

We tested compression by putting a unique text marker into a known PDF, processing it, and then checking the result instead of judging success only by file size.

Question and result

Does NoblePDF's default structural compression preserve real PDF text, or does it make the file smaller by turning pages into pictures? We tested the same qpdf WebAssembly arguments used by the public Compress PDF page against a deliberately uncompressed three-page fixture.

22,656 → 2,800 bytesFresh reproducibility fixture, 87.6% smaller.
3 / 3 markersThe unique text marker remained extractable on all three pages.
0 image XObjectsNo page was replaced with a raster image.

Exact fixture and command

The source fixture contains ordinary Helvetica PDF text with the unique marker NOBLEPDF-COMPRESS-MARKER-58368 once on each page. Its content streams were deliberately left uncompressed so structural recompression had something measurable to change.

--stream-data=compress --object-streams=generate --recompress-flate --compression-level=9
CheckSourcePreserve outputResult
File size22,656 bytes2,800 bytesSmaller without rasterising
Marker extraction3 occurrences3 occurrencesPASS
Pages33PASS
Image XObjects00PASS
Independent parserReadableReadablePASS

Download the exact files

How we checked searchability

We opened both files with an independent PDF parser and extracted text from every page. The exact marker was recovered three times from the source and three times from the compressed output. Page resource dictionaries were also inspected for image XObjects; neither file contained one.

Why this matters: a screenshot-like PDF can look fine while losing text selection, search, copy/paste and vector sharpness. File size alone is therefore not our success criterion.

What this does and does not prove

This fixture demonstrates the design property of NoblePDF's Preserve text & vectors path: it can reduce structural overhead without drawing healthy text pages to a canvas. It does not promise that every PDF will shrink. Already-optimised PDFs, encrypted documents, unusual object structures and image-heavy scans can behave differently.

Reproduce the user-facing check yourself

  1. Download the source fixture above.
  2. Open Compress PDF.
  3. Leave Preserve text & vectors selected.
  4. Compress and download the result.
  5. Search for NOBLEPDF-COMPRESS-MARKER-58368 in the output and compare file size.

Why we used an independent parser

The compressor and the checker should not be the same component. If NoblePDF reported success using only its own rendering code, the test could miss a defect shared by both paths. For this fixture, the output was reopened separately, text was extracted page by page, and the PDF resource dictionaries were inspected for image XObjects.

The source deliberately contains repeated uncompressed text, so the 87.6% size reduction is not a prediction for ordinary documents. It is a controlled demonstration that the qpdf structural path can compress PDF streams and object structure without changing the page into pixels. A modern already-optimised PDF may shrink only a little or not at all.

Hashes make the fixture identifiable

The SHA-256 values printed beside the downloads identify the exact files measured on this page. If the fixture changes in a future NoblePDF build, its hash changes too. That is more useful than publishing a screenshot of a file-size label because the reader can download and inspect the same bytes.

For this run the source hash begins 5599ac21 and the compressed output begins 6620cd25. The complete values are also published in the fixture checksum file.