First confirm it is actually a PDF
A filename ending in .pdf does not guarantee the bytes are a PDF. Downloads can be interrupted, login pages can be saved with the wrong extension, and email gateways can replace files with HTML error responses. A real PDF normally begins with a PDF header near the start of the file and contains PDF objects and trailer/cross-reference information.
Common failure categories
| Symptom | Possible cause |
|---|---|
| File is extremely small | Incomplete download, HTML error page or empty export |
| Viewer asks for a password | Encryption, not corruption |
| One viewer opens it, another refuses | Malformed structure that a tolerant parser repairs automatically |
| Opening stops near the end | Truncation or damaged cross-reference/trailer data |
| Pages render but text extraction fails | Font/encoding issue or image-only pages rather than file corruption |
Cross-reference problems are common in damaged PDFs
Traditional PDFs contain cross-reference information that tells the reader where objects are located. Newer PDFs can store similar information in cross-reference streams. If offsets are wrong or the end of the file is missing, strict readers may fail. Some tools can reconstruct a usable index by scanning for objects, but repair is not guaranteed.
Do not confuse encryption with corruption
A password-protected PDF can be perfectly healthy. If the document requires a password and you do not have it, repeatedly feeding it into converters will not repair anything. Use the legitimate password and an unlock workflow first.
Try a second independent viewer
If one application fails, test another reputable viewer. If two unrelated readers fail, the likelihood of actual file damage is higher. If only a browser tab fails on a very large document, memory pressure or browser resource limits may be the real cause.
Repair can change the file
A repair utility may rebuild cross-reference tables, discard invalid objects or normalise streams. That can save a document, but it may also drop malformed annotations, signatures or metadata. Treat the repaired PDF as a new derived file and compare it with the original.
A careful recovery sequence
- Download the source again if possible.
- Check file size and extension.
- Try a second viewer.
- If encrypted, use the correct password.
- Make a duplicate before repair attempts.
- Use a reputable repair tool if the structure is damaged.
- After repair, compare page count, annotations, signatures and important content.
When to stop trying automated repair
If the PDF is evidence, a signed legal record, an archival master or otherwise high-stakes, aggressive repair may destroy information that a specialist could recover. Preserve the original bytes and work on copies.