A practical PDF glossary
This glossary focuses on terms that affect real document behaviour in NoblePDF: editing, conversion, OCR, redaction, compression, page geometry, privacy and accessibility. It is intentionally more specific than a list of file-format buzzwords.
- AcroForm
- The traditional interactive form system in PDF. It stores fields such as text boxes, checkboxes and buttons separately from ordinary page content.
- Annotation
- An object associated with a page, such as a note, link, highlight or widget. Annotations can be interactive and are not necessarily part of the base page content stream.
- ArtBox
- An optional page box describing the extent of meaningful artwork or content.
- BleedBox
- An optional print-production page box showing the region content may extend into for bleed beyond the final trim.
- Born-digital PDF
- A PDF created directly from software rather than by photographing or scanning paper. It often contains real text and vector objects.
- CropBox
- The rectangle a viewer normally uses as the visible page region. Changing it can hide content without deleting that content. Read the detailed explanation.
- Digital signature
- A cryptographic signature that binds a certificate-backed signature to a byte range of a PDF. It is not the same thing as a picture of a handwritten signature.
- DPI / PPI
- Dots or pixels per inch. For PDF images, effective resolution depends on both image pixel dimensions and the physical size at which the image is placed. Read the detailed explanation.
- Embedded font
- A font program stored inside the PDF so the intended glyphs can be reproduced without relying entirely on fonts installed on the viewer device. Read the detailed explanation.
- Flattening
- Converting interactive or overlay content into ordinary page graphics. Flattening can remove editability and must not be confused with secure redaction.
- Flate compression
- A lossless compression method commonly used for PDF content streams and some image data.
- Glyph
- The visual shape used to display a character. PDF character codes do not always map directly to Unicode characters.
- Incremental update
- A save technique that appends a new PDF revision instead of rewriting the complete file. Earlier object versions may remain in the byte stream. Read the detailed explanation.
- Linearized PDF
- A PDF arranged for progressive display over a network, sometimes called Fast Web View. It is different from incremental saving.
- MediaBox
- The required page rectangle defining the broad physical page canvas.
- Object stream
- A compressed PDF structure that can store multiple indirect objects together to reduce overhead.
- OCR
- Optical character recognition: estimating text from pixels in a scan or image. OCR can add searchability but does not recreate the original source document structure. Read the detailed explanation.
- Output intent
- A colour-management description used by standards such as PDF/A or print workflows to describe the intended output colour condition.
- PDF/A
- A family of ISO-standardised PDF profiles intended for long-term preservation. Conformance requires more than simply changing metadata or a filename. Read the detailed explanation.
- Rasterisation
- Rendering text, vectors and other page objects into pixels. Rasterisation can simplify final output but usually sacrifices selectable text and vector scalability.
- Redaction
- A process intended to permanently remove sensitive information from the delivered document, not merely cover it visually. Read the detailed explanation.
- Searchable scan
- A scanned page image combined with an OCR text layer so search and copy may work while the visible page remains the original scan.
- Subset font
- An embedded font containing only the glyphs needed by the document rather than the entire font program.
- Tagged PDF
- A PDF containing a logical structure tree used to describe headings, paragraphs, lists, tables, figures and reading order for accessibility and reflow. Read the detailed explanation.
- ToUnicode map
- A mapping that helps translate PDF character codes to Unicode, improving text extraction, search and copy.
- TrimBox
- The intended finished dimensions of a printed page after trimming.
- User space
- The PDF page coordinate system in which content positions and page boxes are expressed.
- Vector graphics
- Shapes described mathematically rather than as pixels. Vectors can remain sharp at high zoom and are common for text-like line work, diagrams and logos.
- Whiteout
- A visual covering or replacement area. Unless the workflow removes or irreversibly flattens underlying data, whiteout alone should not be treated as secure redaction.
- XMP
- Extensible Metadata Platform metadata stored in an XML packet, often used for richer document metadata alongside the PDF information dictionary.
- XRef / cross-reference
- Information that helps a PDF reader locate indirect objects. Damaged cross-reference data is a common reason malformed PDFs need repair.
Why these terms matter
Many PDF problems come from treating the file like a screenshot or a word-processing document. Understanding page boxes explains why crops can be reversible. Understanding object streams explains why structural compression can reduce size without rasterising. Understanding OCR explains why a searchable scan may still lack semantic structure. Understanding incremental updates explains why secure redaction needs more than a black rectangle.
Keep learning
The NoblePDF Learning Centre links each concept to a deeper explanation, while PDF Guides cover the workflows most directly connected to NoblePDF tools.