PDF redaction failures that made the news, and how each one happened
By the getPDF team · Published 11 October 2026
The short answer
The best-known PDF redaction failures share 1 cause: black shapes were drawn over text that stayed in the file, so readers selected the black, copied it and pasted the hidden words. It happened to a US military report in 2005, a UK Ministry of Defence safety report in 2011 and a filing in the Manafort case in 2019. The check that catches it takes 30 seconds: select all, copy, paste into a plain text editor, search for what you removed.
The 3 cases, from the reporting at the time
Each case below is told from news reports published when it happened; the source and date are named in each. They are not unusual. They are the ones that were noticed.
May 2005: the US military report on the death of Nicola Calipari
Nicola Calipari, an Italian military intelligence officer, was shot by US troops at a checkpoint in Baghdad in March 2005, shortly after securing the release of a kidnapped Italian journalist. The US military published an unclassified version of its report as a PDF, with large sections blacked out.
The black was only drawn over the text. As The Register reported on 3 May 2005, anyone could copy the text into a word processor and read the blacked-out material, and the full text spread widely online before the PDF was pulled. Washington Technology’s report the same month quoted the military calling its procedures inadequate.
Technical cause: a black overlay on live text; the file was never saved with the covered text removed.
April 2011: the UK Ministry of Defence submarine safety report
In 2009 the Defence Nuclear Safety Regulator wrote advice on choosing the propulsion plant for the planned Successor submarines, marked “Restricted”. After Freedom of Information requests, the Ministry of Defence released it with passages redacted, and it was posted on the UK Parliament’s website in February 2011.
The redacted passages were still in the PDF’s text, so they could be copied and pasted. Google had also made a searchable HTML version of the PDF, which showed the hidden text to anyone who opened it. The Register reported this on 18 April 2011; after the story, the PDF was replaced by a version made of page images.
Technical cause: the same: covered text kept in the file, and a search engine read it as text.
January 2019: the filing in the Manafort case
On 8 January 2019, lawyers for Paul Manafort filed a response to the special counsel’s claim that he had lied to investigators. Several passages were blacked out. As Mother Jones reported that day, each redacted section could be unmasked by highlighting it and pasting it into a text document. The exposed passages included the allegation that he had shared 2016 campaign polling data with an associate, which became a major news story.
Technical cause: black boxes over live text, in a document filed on a public court docket.
The common thread
Every case used the same method: draw something black over the text, export or save, publish. That method fails because of how a PDF stores a page. The page is a list of drawing instructions. Text is an instruction that says “these characters, in this font, here”. A black box is a later instruction that paints over them. Everything that reads text (copy, search, screen readers, search engines, text extraction) reads the text instruction and never looks at what was painted over it.
The format makes the mistake easy and the fix cheap. Real redaction takes the characters out of the text instruction, paints out the pixels of any picture under the box, and only then draws the black. Redact a PDF properly shows what that looks like in practice, measured on our invoice test file.
The 30-second check that would have caught all 3
- Open the redacted PDF in any viewer.
- Press Ctrl+A (Cmd+A on a Mac), then Ctrl+C.
- Paste into Notepad, TextEdit or any plain text editor.
- Search for a word or number you redacted.
If it is there, the redaction failed, whatever the page looks like. For a long file, drop it on PDF to Text instead: it extracts every character on every page, including invisible text that a viewer will not let you select, and you search the result.
Two more places to look, because they are outside the pages: the document properties (File, Properties in most viewers, or Inspect), where a title or author can repeat what the page hides, and the file name.
What these cases are not
They are not evidence that PDFs are unsafe or that redaction does not work. Each was a process gap: the step that removes the text was skipped, and nobody ran the check before publishing. The 2011 fix, page images, works but makes the whole document unsearchable and inaccessible to screen readers. True redaction removes only what is under the boxes and keeps the rest of the document as text.
The honest part
Tools fail too, so verify the output, not the tool’s promise. getPDF’s redaction is tested by decompressing every stream of the saved file and searching it for the removed text, and the test fails if one stale copy remains. That is a reason to trust it more than a black rectangle, not a reason to skip the 30-second check, which costs nothing.
Two limits that the check does not cover. Copy and paste only proves the text is gone; on a scanned page the words are pixels, so look at the page itself, zoomed in. And no check finds what the boxes left visible: a name in a footer, the same number on page 12, or a sentence that makes the hidden one obvious. Read the redacted file as a stranger would before it goes out.
Questions
Why can you copy text from under a black box in a PDF?
Because the box is a separate drawing on top of the page. The text instruction underneath is unchanged, and copy, search and text extraction read text instructions, not what the screen shows.
Does printing a redacted PDF to a new PDF fix it?
Often, but not reliably. Many print-to-PDF drivers keep text as text, so the covered words survive. Saving the pages as pictures does remove all text, and true redaction removes only what is under the boxes and keeps the rest searchable.
How do I check a redacted PDF in 30 seconds?
Open it, select all (Ctrl+A or Cmd+A), copy, paste into Notepad or TextEdit, and search for a word you redacted. If it is there, the redaction failed. PDF to Text runs the same check on every page, invisible text included.
Were the people involved using bad software?
The reports describe the same method each time: black shapes or highlighting over live text, then an export to PDF. The tools did what they were told; the step that removes the text was never run.
The tools for this job
Keep reading
- Redact a PDF properly: remove the text, never just cover itA black rectangle hides nothing: the text underneath still copies out.
- Redact a scanned PDF: remove the pixels, not just add a boxA scan is a picture, and a drawn black box is a second object anyone can delete.
- Clean a PDF before publishing it: the 7-step pass for hidden leftoversMetadata, comments, bookmarks, attachments, old form data, earlier versions: the 7-step pass that catches what a visual check misses before a PDF goes public.
- What the metadata in your PDFs reveals, and how to remove itAuthor names, employers, software versions, edit dates: the metadata a typical PDF carries, how to read yours in 10 seconds, and how to remove it all.
- PDF privacy and protection: the complete guideWhat a PDF password really protects, how true redaction works, what metadata leaks, and how to send files safely.