Skip to content

PDF redaction failures that made the news, and how each one happened

By the getPDF team · Published 11 October 2026

The short answer

The best-known PDF redaction failures share 1 cause: black shapes were drawn over text that stayed in the file, so readers selected the black, copied it and pasted the hidden words. It happened to a US military report in 2005, a UK Ministry of Defence safety report in 2011 and a filing in the Manafort case in 2019. The check that catches it takes 30 seconds: select all, copy, paste into a plain text editor, search for what you removed.

Redact PDFRemoves content, not just a black box. Free, runs on your device.

The 3 cases, from the reporting at the time

Each case below is told from news reports published when it happened; the source and date are named in each. They are not unusual. They are the ones that were noticed.

A timeline from 2005 to 2019 with three marks: May 2005, the US military report on the Calipari shooting; April 2011, the UK submarine safety report; January 2019, the Manafort court filing. Under each, the cause: black boxes over live text.May 2005April 2011January 2019US military reporton the Calipari shootingUK MoD submarinesafety reportManafort defencefilingblack boxes, textleft underneathtext kept in the file,also indexed by Googleselect, copy, pasteread every passage
3 documented redaction failures, 2005 to 2019. Each one covered text instead of removing it.

May 2005: the US military report on the death of Nicola Calipari

Nicola Calipari, an Italian military intelligence officer, was shot by US troops at a checkpoint in Baghdad in March 2005, shortly after securing the release of a kidnapped Italian journalist. The US military published an unclassified version of its report as a PDF, with large sections blacked out.

The black was only drawn over the text. As The Register reported on 3 May 2005, anyone could copy the text into a word processor and read the blacked-out material, and the full text spread widely online before the PDF was pulled. Washington Technology’s report the same month quoted the military calling its procedures inadequate.

Technical cause: a black overlay on live text; the file was never saved with the covered text removed.

April 2011: the UK Ministry of Defence submarine safety report

In 2009 the Defence Nuclear Safety Regulator wrote advice on choosing the propulsion plant for the planned Successor submarines, marked “Restricted”. After Freedom of Information requests, the Ministry of Defence released it with passages redacted, and it was posted on the UK Parliament’s website in February 2011.

The redacted passages were still in the PDF’s text, so they could be copied and pasted. Google had also made a searchable HTML version of the PDF, which showed the hidden text to anyone who opened it. The Register reported this on 18 April 2011; after the story, the PDF was replaced by a version made of page images.

Technical cause: the same: covered text kept in the file, and a search engine read it as text.

January 2019: the filing in the Manafort case

On 8 January 2019, lawyers for Paul Manafort filed a response to the special counsel’s claim that he had lied to investigators. Several passages were blacked out. As Mother Jones reported that day, each redacted section could be unmasked by highlighting it and pasting it into a text document. The exposed passages included the allegation that he had shared 2016 campaign polling data with an associate, which became a major news story.

Technical cause: black boxes over live text, in a document filed on a public court docket.

The common thread

Every case used the same method: draw something black over the text, export or save, publish. That method fails because of how a PDF stores a page. The page is a list of drawing instructions. Text is an instruction that says “these characters, in this font, here”. A black box is a later instruction that paints over them. Everything that reads text (copy, search, screen readers, search engines, text extraction) reads the text instruction and never looks at what was painted over it.

The format makes the mistake easy and the fix cheap. Real redaction takes the characters out of the text instruction, paints out the pixels of any picture under the box, and only then draws the black. Redact a PDF properly shows what that looks like in practice, measured on our invoice test file.

The 30-second check that would have caught all 3

  1. Open the redacted PDF in any viewer.
  2. Press Ctrl+A (Cmd+A on a Mac), then Ctrl+C.
  3. Paste into Notepad, TextEdit or any plain text editor.
  4. Search for a word or number you redacted.

If it is there, the redaction failed, whatever the page looks like. For a long file, drop it on PDF to Text instead: it extracts every character on every page, including invisible text that a viewer will not let you select, and you search the result.

Two more places to look, because they are outside the pages: the document properties (File, Properties in most viewers, or Inspect), where a title or author can repeat what the page hides, and the file name.

What these cases are not

They are not evidence that PDFs are unsafe or that redaction does not work. Each was a process gap: the step that removes the text was skipped, and nobody ran the check before publishing. The 2011 fix, page images, works but makes the whole document unsearchable and inaccessible to screen readers. True redaction removes only what is under the boxes and keeps the rest of the document as text.

The honest part

Tools fail too, so verify the output, not the tool’s promise. getPDF’s redaction is tested by decompressing every stream of the saved file and searching it for the removed text, and the test fails if one stale copy remains. That is a reason to trust it more than a black rectangle, not a reason to skip the 30-second check, which costs nothing.

Two limits that the check does not cover. Copy and paste only proves the text is gone; on a scanned page the words are pixels, so look at the page itself, zoomed in. And no check finds what the boxes left visible: a name in a footer, the same number on page 12, or a sentence that makes the hidden one obvious. Read the redacted file as a stranger would before it goes out.

Questions

Why can you copy text from under a black box in a PDF?

Because the box is a separate drawing on top of the page. The text instruction underneath is unchanged, and copy, search and text extraction read text instructions, not what the screen shows.

Does printing a redacted PDF to a new PDF fix it?

Often, but not reliably. Many print-to-PDF drivers keep text as text, so the covered words survive. Saving the pages as pictures does remove all text, and true redaction removes only what is under the boxes and keeps the rest searchable.

How do I check a redacted PDF in 30 seconds?

Open it, select all (Ctrl+A or Cmd+A), copy, paste into Notepad or TextEdit, and search for a word you redacted. If it is there, the redaction failed. PDF to Text runs the same check on every page, invisible text included.

Were the people involved using bad software?

The reports describe the same method each time: black shapes or highlighting over live text, then an export to PDF. The tools did what they were told; the step that removes the text was never run.

The tools for this job