Start by asking what kind of document it is
A 200-page text report and a 20-page scanned brochure can have the same file size for completely different reasons. Compression works best when it targets the dominant resource rather than applying the same transformation to every page.
Images are often the biggest contributor
Photographs and scans can contain millions of pixels per page. Multiple high-resolution colour images quickly dominate file size, especially when a document was built for print but is being shared only on screen. Downsampling or recompressing those images can create large savings, but the quality trade-off is visible and sometimes irreversible.
Fonts can matter more than expected
A PDF can embed one or more font programs. Subsetting usually keeps only the glyphs that are used, but badly produced documents may embed whole fonts repeatedly. Documents containing many writing systems or multiple font families can legitimately need more font data.
Duplicate resources waste space
Two pages may use the same logo but store separate copies rather than referencing one shared image object. Similar duplication can occur with fonts, colour profiles and other resources. Structural optimisation can sometimes consolidate or compress these objects without changing page appearance.
Streams may already be compressed
PDF content streams can use compression filters such as Flate. JPEG images are already compressed with DCT. Recompressing already-compressed data may save very little. That is why a structural compression pass can produce almost no change on a well-optimised born-digital PDF.
Embedded files and attachments can dwarf the pages
A PDF can contain arbitrary embedded files. A visually simple PDF could therefore be unexpectedly large because it carries attachments. Before destroying page quality to chase a size target, check whether the extra bytes are actually part of the visible pages.
| Likely source of size | Best first question |
|---|---|
| Scanned pages | Can image resolution/colour depth be reduced while keeping text legible? |
| Photographs | Are images much larger than their displayed size requires? |
| Text/vector report | Can streams/object structures be optimised without rasterising? |
| Many fonts | Are fonts subset and shared appropriately? |
| Attachments | Are embedded files actually needed in the delivered PDF? |
Why “smaller” is not the only success metric
Turning every page into JPEG may cut size while destroying selectable text, links, accessibility structure and vector sharpness. A 30% smaller file can be a worse document. NoblePDF therefore defaults its Compress PDF tool to Preserve text & vectors and labels image-flattening modes separately.
A diagnostic workflow
- Check whether pages are text or images.
- Try structural optimisation first.
- If the result barely changes and the PDF is image-heavy, test an image-based mode on a copy.
- Compare file size, legibility and searchability.
- Keep the original if the compressed copy does not provide a meaningful benefit.