Skip to content

Redact a scanned PDF: remove the pixels, not just add a box

By the getPDF team · Published 11 October 2026

The short answer

On a scanned page the words are pixels in a picture, so redaction has to change the picture itself: a black box drawn on top is a separate object that anyone can delete. Open the scan in Redact PDF, drag a box over each area (on a scan that went through OCR, Find personal data can mark the numbers for you), and save: getPDF paints those pixels black inside the picture, removes any invisible OCR text under the box, and writes a new file. Then zoom in on the saved copy and run it through PDF to Text to check.

Redact PDFRemoves content, not just a black box. Free, runs on your device.

Why scans fail differently

A scanned page is 1 picture, usually a JPEG, filling the page. There is no text in it, so the copy-and-paste failure that exposes covered text does not apply. Two other things do.

The box is a separate object. Many apps “redact” a scan by putting a black rectangle on top of it: a comment, a shape, or a drawing in the page. The picture underneath is untouched. Open the file in any editor, select the rectangle, delete it, and the scan shows the number again. Extract the images from the PDF and the full picture comes out with no box at all.

The OCR text layer. A scan that went through OCR (many scanner apps do it automatically) carries an invisible text layer: every recognised word, placed over the picture so you can search and copy. Paint the pixels black and leave that layer, and select-all still pastes the account number.

Two scanned receipts. On the left, a black rectangle floats above the picture, and the card number is still printed in the picture below it. On the right, the card number area inside the picture itself is black and nothing floats above it.Box on top4111 1111 1111 1111a second object: delete it and the number showsPixels paintedthe picture itself is black there
Left: a black box on top of a scan. Delete the box, or extract the picture, and the number is back. Right: the pixels themselves are painted black, so the picture in the file no longer holds the number.

What getPDF does to a scan under a box

We measured it on our scanned lease test file: 1 A5 landscape page, a greyscale scan at 200 dpi, 1,165 by 826 pixels, stored as a 34.8 KB JPEG.

  • A box over part of the scan: the picture stays 1,165 by 826 pixels, and the covered pixels are black in the picture itself. We removed the black box from the saved page to look underneath: where the page had been paper white, the picture was now black. The JPEG is written again once, at quality 90, so the file went from 34.8 KB to 45.3 KB. It took about 0.1 seconds.
  • A box over the whole scan: the picture is removed from the file. The saved file was 895 bytes: a page with a black box and nothing else.
  • An OCR’d scan: after OCR, a box over the rent amount took out 4 invisible characters and painted the picture under them. The text layer now reads “Monthly rent: EUR”.
  • Redact first, OCR after: we redacted the account number line of a scanned bank statement, then ran OCR on the saved page. OCR read the black area as a few letters of noise; the number did not come back.

How to redact a scanned PDF

  1. Open the scan at /redact-pdf. The Redact tool is picked. Nothing is uploaded.
  2. Zoom in on the area so you can see the edges of the characters.
  3. Drag a box over each thing that must go, with a margin of a few pixels around the ink. Scans are often slightly tilted, so a box that fits the first digit can miss the last one. If the scan went through OCR, click Find personal data first to mark the emails, phone, card and account numbers, then add boxes for the rest.
  4. Look for repeats. A receipt prints the card number twice, a statement repeats the account number on every page, a form has a name in the header and again by the signature.
  5. Click Save. The status line says how many areas, characters and pictures were taken out; on a plain scan it counts 1 picture for each scanned page you boxed. The saved file is named -redacted.

Can the finder mark a scan for me?

Yes, once the scan has a text layer. Find personal data reads text, so on a scan without OCR it finds nothing: run OCR PDF first, then open the result here. On an OCR’d scan it reads the invisible layer and boxes each hit at the height of the printed line, so the box covers the ink, not only the invisible words.

We measured it on 11 October 2026 on a scanned bank statement page, 11-point type at 200 dpi. OCR read it at 94 % confidence. The finder listed both IBANs, the card number, the email and the mobile number, with boxes 10 to 13 points tall over the printed lines. We marked them and saved; the status line counted 1 picture painted. A second OCR of the redacted page read none of the numbers back.

Still look at every box before you save. OCR can misread a digit, and a misread number fails its checksum or its shape and is not listed. And a box sits on what OCR saw, so it can clip a letter of the next word. Add a box by hand wherever the finder’s box looks short.

Check the saved file

  1. Zoom in to 400 % on each redacted area of the -redacted file. You should see solid black with clean edges and no grey ghost of the digits.
  2. Run PDF to Text on it. On a scan with an OCR layer, search for the removed number. On a scan without one, the result is empty, which is expected.
  3. Delete the box to be sure, if you want proof: open the saved file in the editor, select the black box with the Select tool and delete it. Underneath, the picture itself is black. (Do not save that copy.)

The honest part

  • Check the neighbourhood. A tight box misses a crease, a stamp or a smudged copy of the same number a few lines down. A carbon copy or a second page printed with the same details survives a box on page 1.
  • Pixels outside the box stay as they were, apart from the one JPEG re-encode at quality 90. That is fine for reading, but if you need the scan archived bit for bit, keep the original separately.
  • A scan is also metadata. Scanner apps often write their name and the scan date into the file properties. Redaction does not touch them; Remove PDF metadata does.
  • OCR after redaction is fine; OCR on a different copy is not. If someone else OCRs the original scan, their text layer has everything. Send only the -redacted file.

For a digital document (text you can select), Redact a PDF properly covers the rest, and Scanned PDF to searchable explains the OCR text layer in detail.

Questions

Can a black box over a scanned page be removed?

If it is a box drawn on top, yes. It is a separate object (often a comment or a shape), and any PDF editor can select and delete it, which shows the scan underneath. Real redaction paints the pixels of the scan itself black.

Does getPDF find personal data on a scan?

Yes, after OCR. The finder reads text, so a scan without OCR gives no hits. On an OCR'd scan it boxes each hit at the printed line's height, and saving paints those pixels black. Check the boxes before saving: a number OCR misread is not found.

Should I redact before or after OCR?

In getPDF either order works: redaction removes the invisible OCR words under a box together with the pixels. In other tools, redact first and OCR after, so no text layer holds what you covered.

Does redaction make my scan blurry?

No visible change outside the boxes. A JPEG scan is written again once at quality 90, so the file grows a little: our 1-page lease scan went from 34.8 KB to 45.3 KB.

The tools for this job