PDF internals

Incremental PDF Updates: Why Old Data Can Remain in a File

A PDF does not always get rewritten from the first byte to the last when you save it. Some applications append a new revision instead.

Why incremental saving exists

PDF files can be large, and certain workflows benefit from preserving prior byte ranges. An incremental update appends changed or new objects plus a new cross-reference section and trailer information. The viewer follows the newest revision to determine the current state.

The old bytes may still be there

If an object is superseded by a newer version, the previous object can remain earlier in the file. It may no longer be part of the visible current document, but forensic tools can sometimes inspect earlier revisions. This is useful for signatures and document history, but dangerous if someone believes “I changed the visible page, therefore the old data is gone.”

Why this matters for redaction

Drawing a black rectangle over text and saving incrementally can leave the original text object completely intact underneath. Even removing an object from the current page tree may not be sufficient if a previous revision remains inside the file. Secure redaction therefore requires a workflow designed to eliminate the sensitive content from the delivered result.

Practical rule: for sensitive redaction, verify the exported file - not just the screen. Search it, copy from it, and when the stakes justify it, inspect the decompressed PDF objects with an independent tool.

Digital signatures are one reason revisions matter

Cryptographic PDF signatures sign a specific byte range. Later permitted changes can be appended as incremental revisions so the original signed bytes remain untouched. A validator can then reason about what was signed and what changed afterward.

Linearisation is a different concept

A “fast web view” or linearised PDF is arranged so a viewer can begin displaying pages before the entire file downloads. That has nothing to do with incremental revisions, although both features affect the physical ordering of objects inside the file.

How to reduce historical-data surprises

NoblePDF's destructive-page approach

When the editor detects destructive edits such as redaction, whiteout replacement or existing-text replacement, that page is exported through a flattened path rather than preserving the original page objects underneath. Non-destructive annotations use a different path so ordinary text and vectors can remain selectable. Separating those cases is important because “maximum fidelity” and “remove the underlying content” are sometimes conflicting goals.