Fix a broken PDF: find the real fault, recover what you can
By the getPDF team · Published 11 October 2026
The short answer
Most “corrupted” PDFs are not deeply damaged: the download was cut short, or the file is not really a PDF at all. Re-download first and compare the byte count with the original; a size mismatch is the proof. Then drop it on the inspector below: if it opens, you see its pages, fonts and whether the pages are scans; if it “could not be read as a PDF”, the file is cut short, damaged or not a PDF at all. Everything runs on your device; nothing is uploaded.
Start with a look inside the file
Before any fix, find out what you actually have. Drop the file here: the inspector reads it on your device and shows the PDF version, the page count, every font and whether it is embedded, the images and their compression, how many pages are scans, the metadata, and what makes the file big. If the engine cannot open the file at all, it says the file “could not be read as a PDF. It may be damaged, or not a PDF at all”, which is the answer to the first question below.
Try it here, nothing is uploaded
That 1 look settles the 2 questions that decide everything after. Is this file a readable PDF at all? And is the text real text or a photograph of text? A scan and a damaged file produce similar complaints (cannot search, cannot select, odd behaviour in some viewers) but need opposite fixes.
One thing to say plainly up front: getPDF does not have a repair tool yet. It is coming. Until it ships, this page keeps the manual routes current, and they fix the large majority of cases because the large majority of cases are not true corruption.
The symptom map
Work from what you see, not from the scariest error message. The symptoms below are ordered by how often they turn up, and each gets its own section with the causes ranked.
| What you see | Most likely cause | First move |
|---|---|---|
| Will not open, “damaged” error | Cut-off download | Re-download, compare file sizes |
| Opens, but text is boxes or gibberish | Fonts not embedded | Check the font list in /inspect |
| Cannot select or search text | It is a scan | OCR, not repair |
| Prints wrong or blank | Annotation layers or scaling | Try a second viewer, check print settings |
| Links stopped working | Conversion dropped annotations | Re-export from the source |
| File got huge after saving | Pages were saved as images | Compress it, re-export if possible |
| Blank pages in 1 viewer, fine in another | Rendering differences | Use the viewer that works, make a flat copy |
The file will not open
The hard “File is damaged and could not be repaired” moment. 4 causes, most likely first.
1. The download or transfer was cut short. This is the most common cause by a wide margin. A PDF keeps its table of contents (the cross reference table) at the end of the file, so a transfer that stops at 90 percent does not give you 90 percent of a working document; it gives you a file whose index is missing, and strict viewers refuse it outright. The test takes 30 seconds: check the file’s size in its properties and compare it with the size at the source (the download page, the email, the sender’s copy). If the numbers differ, the file is truncated, and the fix is not repair, it is a fresh download over a stable connection. Most “corrupt” files end here. (We cut a 5.7 MB test file at 90 % of its length: the inspector could not read it either.)
A file that is simply too big for the browser gets a different message: it “is too big for this browser to open”, with how much memory opening it needs and how much the browser has. That file is not damaged; a computer with more memory, or a desktop browser instead of a phone, opens it.
2. It is not a PDF. Files get renamed in transit: a Word document, an image or an HTML error page saved with a .pdf extension. A viewer asked to open one reports damage, because as a PDF it is damaged, and /inspect says it could not be read as a PDF. To tell which case you have, open the file in a plain text editor (Notepad on Windows, TextEdit on a Mac) and look at the first characters: a real PDF begins with %PDF-, and a cut-off one still does. If yours begins with anything else (PK means a ZIP or Office file, an HTML tag means a web page saved by mistake), rename it to the right extension or go back to the source and download the actual document.
3. True corruption. Real byte damage happens: a failing disk, a USB stick pulled mid-write, a mail system that mangled the attachment. Try the file in another viewer before giving up, because strictness varies enormously. Acrobat is the strictest of the common viewers; Chrome, Edge and Preview tolerate a lot, and a file that Acrobat rejects often renders fully in a browser tab. If any viewer shows the pages, print to PDF from that viewer: the rescue copy loses interactive parts (form fields, links) but keeps the look, and it opens everywhere. For technical readers, the command line tool qpdf can rebuild a broken cross reference table and often restores a refused file to full health.
4. The viewer is too old for the file. A PDF 2.0 file with modern encryption can refuse to open in software from the 2000s. The version sits in the file header and /inspect shows it; the fix is any current viewer, and the browser you are reading this in is one.
If all 4 fail, ask for a fresh copy from whoever made the file. It is not an admission of defeat; it is usually the fastest route, and the only one when the damage is real.
What corruption actually is
A PDF is not a stream of pages. It is a bag of numbered objects (pages, fonts, images, metadata) plus a cross reference table at the end that records where each object sits, and a trailer that points at the table. The practical consequence: where the damage lands matters more than how big it is.
A damaged or missing table stops strict viewers cold even though every page may still be in the file, which is why rebuilding it (a browser does this silently, qpdf does it on demand) recovers so many “hopeless” files. Damage in the middle of the file is the opposite case: the file opens, but 1 image is grey from halfway down or 1 page fails, and the loss is real but local.
Text shows as boxes or gibberish
The file opens and the layout looks right, but the words are squares, dots or the wrong characters. This is not corruption; it is fonts.
Boxes mean missing glyphs. The font was not embedded in the file, your machine does not have it, and the viewer has nothing to draw. Open the file in /inspect and read the font list: each font says embedded or not. A document with non-embedded fonts displays differently on every machine it visits. The quick fix is trying another viewer, because each substitutes missing fonts differently and one may pick a font you have. The real fix is upstream: whoever makes the file re-exports it with fonts embedded.
Gibberish means broken encoding. The bytes in the file no longer map to the right characters, so the viewer draws letters, just the wrong ones. This usually cannot be fixed inside the file. But if the page renders correctly on screen and only copying or searching produces junk, the words are recoverable: run the file through OCR, which reads the rendered page the way your eyes do and rebuilds a clean text layer under it.
You cannot select or search the text
Nothing is broken. If dragging across a paragraph selects nothing and Ctrl+F finds nothing, the page is almost certainly a picture of text: a scan or a photo placed on the page, with no text objects behind it. /inspect confirms it in 1 glance; a scan shows 1 large image per page and little or no text content.
The fix is OCR, not repair. The complete guide is make a scanned PDF searchable: it reads the pages on your device and writes an invisible text layer under the image, so the document looks the same but selects, searches and copies like a born-digital file.
It prints wrong or blank
A file that looks right on screen and wrong on paper is usually a settings or layers problem, not a broken file.
- Annotations do not print. Highlights, form field contents and signatures often live on a separate layer that print dialogs skip by default. Look for a “document and markups” or “print with annotations” option in the print dialog.
- Scaling shifts everything. An A4 page printed on Letter paper (or the reverse) with “fit to page” off gets cropped; with it on, margins drift. Check the size setting matches the paper in the tray.
- Font substitution strikes again. A non-embedded font can print differently than it displays, because the printer driver substitutes on its own. The font list in /inspect predicts this before paper is wasted.
- A blank or failed page usually means 1 element the printer chokes on. Printing from a different viewer, or using the “print as image” fallback, gets the job out; the page prints slower and text softens slightly, but it matches the screen.
Links stopped working after conversion
A link in a PDF is not the blue underlined text; it is an invisible clickable rectangle stored separately, above the page. That explains the whole symptom: any process that redraws pages without carrying those rectangles over (print to PDF is the classic, since printing has no concept of a hyperlink) produces a file where the text survives and every link is dead.
There is nothing to repair afterwards, because the links were never written into the new file. The fix is the route, not the file: export from the source application (Word’s save as PDF, a browser’s HTML export) rather than printing, and the links travel. Links that point at local files break for a different reason: they are relative, and die the moment the PDF leaves the folder structure they pointed into.
The file ballooned after saving
A 2 MB document goes through 1 edit or save and comes out at 40 MB. The usual cause: the application rasterized the pages, replacing crisp, tiny drawing instructions with a full-page photograph of each page. Other culprits are a scan at needless resolution and editors that append every old version to the end of the file instead of rewriting it.
/inspect shows exactly where the bytes sit: when the size anatomy is nearly all images on a document that used to be text, the save rasterized it. Compressing the file on your device gets the size back down, and the before and after is shown. What compression cannot restore is the lost vector crispness; only re-exporting from the source document brings that back. The deep dive on size, with the levers and the real numbers, is make a PDF smaller.
Pages blank in one viewer, fine in another
A PDF is a set of drawing instructions, and at least 5 major engines interpret them: Acrobat’s, PDFium in Chrome and Edge, Apple’s in Preview, Firefox’s pdf.js, and the long tail. They disagree at the edges. Transparency effects from design tools, optional content layers that 1 viewer hides by default, and oversized page images that a low-memory device fails to decode all produce the same sight: a page blank here, fine there.
First establish which case you have. Open the same file in 2 viewers (a browser and a desktop viewer is the easy pair). Blank in 1 place is a rendering difference, and the file is healthy; blank everywhere, with /inspect showing no content on the page, means the page was exported empty and there is nothing to recover. For the rendering case, the universal fix is a flat copy: print to PDF from the viewer that shows the page correctly. The copy trades editability for a file that displays the same everywhere.
A password is not corruption
A file that demands a password, refuses to print or will not let you copy text is not damaged; it is restricted, and no repair route applies. If you know the password, the unlock tool removes it on your device so the file stops asking. We do not crack passwords: a file encrypted with a password nobody knows stays closed, and honest software says so. The full picture of passwords, permissions and what they do and do not protect is in PDF privacy and protection.
What cannot be recovered, plainly
The honest limits, so you do not spend an evening on a file that is gone:
- Bytes that are missing are missing. A file truncated to half its size lost the second half; nothing reconstructs data that was never saved to disk. What a rescue recovers is the content up to the cut.
- Overwritten bytes are gone too. When disk corruption replaces part of a file, the affected objects (an image, a page’s content) cannot be rebuilt from what remains. Recovery here means extracting the intact parts.
- Encrypted files without the password stay closed. The encryption in modern PDFs is strong by design. Anyone claiming to open them without the key is overstating.
- Nothing on this page uploads your file. Every route here (the inspector, OCR, compression, unlocking) runs in your browser, on your device. You can prove it: run a tool once, so its engine is in your browser, then turn Wi-Fi off, and the tool still works. That matters most exactly when the broken file is a contract or a medical record.
A repair tool for getPDF is on the roadmap: one that rebuilds cross reference tables and salvages intact objects, on your device, like everything else here. Until it ships, the sequence that solves most real cases costs nothing and needs no install: fresh download, size check, /inspect, second viewer, print to a rescue copy.
Prevent the next broken file
Three habits cover nearly everything. Verify large transfers by size: a byte count that matches the source is a transfer that arrived whole. Eject USB drives properly and let cloud folders finish syncing before you close the laptop; most true corruption is an interrupted write. And keep the source document (the Word file, the design file) that made the PDF, because the cleanest repair of all is exporting a fresh one.
Questions
Can getPDF repair a corrupted PDF?
Not yet; a repair tool is on the roadmap. Today the inspector tells you what is wrong, and this guide gives the manual routes that work most often, starting with a fresh download. No step uploads your file; everything runs in your browser.
Why does my PDF open in Chrome but not in Acrobat?
Viewers differ in strictness. Acrobat refuses a file whose internal table it cannot read, while Chrome and Edge rebuild that table and render anyway. A file that opens anywhere is mostly intact; print it to a new PDF from the viewer that works.
Is a password error the same as corruption?
No. A file that asks for a password is healthy; the viewer just will not show it without the key. If you know the password, the unlock tool removes it on your device. Nothing here cracks passwords you do not know.
Can an online PDF repair service fix what this page cannot?
No tool restores bytes that are missing or overwritten; a service that claims to is guessing. The difference that matters is that an upload service receives your document. The routes on this page run on your device, provable with Wi-Fi off.
My file is half the size it should be. Is anything recoverable?
Content up to the cut usually is; everything past it is gone. A forgiving viewer may render the surviving pages, and printing those to a new PDF saves what remains. Compare byte counts with the source to confirm the file was truncated.