Reference

PDF Glossary: 30+ Terms Explained in Plain English

PDF terminology is full of words that sound simple but have precise consequences for editing, privacy and conversion.

A practical PDF glossary

This glossary focuses on terms that affect real document behaviour in NoblePDF: editing, conversion, OCR, redaction, compression, page geometry, privacy and accessibility. It is intentionally more specific than a list of file-format buzzwords.

AcroForm
The traditional interactive form system in PDF. It stores fields such as text boxes, checkboxes and buttons separately from ordinary page content.
Annotation
An object associated with a page, such as a note, link, highlight or widget. Annotations can be interactive and are not necessarily part of the base page content stream.
ArtBox
An optional page box describing the extent of meaningful artwork or content.
BleedBox
An optional print-production page box showing the region content may extend into for bleed beyond the final trim.
Born-digital PDF
A PDF created directly from software rather than by photographing or scanning paper. It often contains real text and vector objects.
CropBox
The rectangle a viewer normally uses as the visible page region. Changing it can hide content without deleting that content. Read the detailed explanation.
Digital signature
A cryptographic signature that binds a certificate-backed signature to a byte range of a PDF. It is not the same thing as a picture of a handwritten signature.
DPI / PPI
Dots or pixels per inch. For PDF images, effective resolution depends on both image pixel dimensions and the physical size at which the image is placed. Read the detailed explanation.
Embedded font
A font program stored inside the PDF so the intended glyphs can be reproduced without relying entirely on fonts installed on the viewer device. Read the detailed explanation.
Flattening
Converting interactive or overlay content into ordinary page graphics. Flattening can remove editability and must not be confused with secure redaction.
Flate compression
A lossless compression method commonly used for PDF content streams and some image data.
Glyph
The visual shape used to display a character. PDF character codes do not always map directly to Unicode characters.
Incremental update
A save technique that appends a new PDF revision instead of rewriting the complete file. Earlier object versions may remain in the byte stream. Read the detailed explanation.
Linearized PDF
A PDF arranged for progressive display over a network, sometimes called Fast Web View. It is different from incremental saving.
MediaBox
The required page rectangle defining the broad physical page canvas.
Object stream
A compressed PDF structure that can store multiple indirect objects together to reduce overhead.
OCR
Optical character recognition: estimating text from pixels in a scan or image. OCR can add searchability but does not recreate the original source document structure. Read the detailed explanation.
Output intent
A colour-management description used by standards such as PDF/A or print workflows to describe the intended output colour condition.
PDF/A
A family of ISO-standardised PDF profiles intended for long-term preservation. Conformance requires more than simply changing metadata or a filename. Read the detailed explanation.
Rasterisation
Rendering text, vectors and other page objects into pixels. Rasterisation can simplify final output but usually sacrifices selectable text and vector scalability.
Redaction
A process intended to permanently remove sensitive information from the delivered document, not merely cover it visually. Read the detailed explanation.
Searchable scan
A scanned page image combined with an OCR text layer so search and copy may work while the visible page remains the original scan.
Subset font
An embedded font containing only the glyphs needed by the document rather than the entire font program.
Tagged PDF
A PDF containing a logical structure tree used to describe headings, paragraphs, lists, tables, figures and reading order for accessibility and reflow. Read the detailed explanation.
ToUnicode map
A mapping that helps translate PDF character codes to Unicode, improving text extraction, search and copy.
TrimBox
The intended finished dimensions of a printed page after trimming.
User space
The PDF page coordinate system in which content positions and page boxes are expressed.
Vector graphics
Shapes described mathematically rather than as pixels. Vectors can remain sharp at high zoom and are common for text-like line work, diagrams and logos.
Whiteout
A visual covering or replacement area. Unless the workflow removes or irreversibly flattens underlying data, whiteout alone should not be treated as secure redaction.
XMP
Extensible Metadata Platform metadata stored in an XML packet, often used for richer document metadata alongside the PDF information dictionary.
XRef / cross-reference
Information that helps a PDF reader locate indirect objects. Damaged cross-reference data is a common reason malformed PDFs need repair.

Why these terms matter

Many PDF problems come from treating the file like a screenshot or a word-processing document. Understanding page boxes explains why crops can be reversible. Understanding object streams explains why structural compression can reduce size without rasterising. Understanding OCR explains why a searchable scan may still lack semantic structure. Understanding incremental updates explains why secure redaction needs more than a black rectangle.

Keep learning

The NoblePDF Learning Centre links each concept to a deeper explanation, while PDF Guides cover the workflows most directly connected to NoblePDF tools.