Skip to content

Find personal data in a PDF before you share it

By the getPDF team · Published 11 October 2026

The short answer

Open the PDF in Redact PDF and click Find personal data: it looks through every page on your device and lists email addresses, phone numbers, IBANs, card numbers, US Social Security, UK National Insurance and EU VAT numbers, and dates, in the page text and in filled form fields. Untick what may stay, click the Mark button, check the boxes and save; the marked text is taken out of the file. Names and street addresses are not found by rules, and the document’s properties, bookmarks and attachments are not searched, so check those yourself.

Run the finder

Redact PDFRemoves content, not just a black box. Free, runs on your device.
  1. Drop the PDF on the page above, or open /redact-pdf. The editor opens with the Redact tool picked.
  2. Click Find personal data in the bar above the page. The panel shows “Looking through page 3 of 40…” while it reads.
  3. Read the list. Hits are grouped by kind, such as “Bank accounts (IBAN)”, with a count like “40 of 40” and the page of each hit.
  4. Untick anything that may stay. Dates start unticked, because most documents are full of harmless ones.
  5. Click Mark 9 for redaction (the number is what you left ticked). Each hit gets an ordinary redaction box on the page.
  6. Scroll through and look at the boxes. Add your own with the Redact tool (press R and drag) for anything the rules cannot find.
  7. Save with Ctrl+S. The status line says what was taken out, for example “Saved with 9 areas redacted: 138 characters, 1 form field taken out”.

Nothing is removed until you save, and every mark undoes. The file downloads with “-redacted” in its name; your original stays as it was.

What it finds, and how it checks

The finder uses rules, not guesses. Each kind has a shape, and the kinds with a checksum are checked, which keeps false hits down:

Kind Example it finds How it is checked
Email addresses jane.cooper@example.com The shape name, at sign, domain
Phone numbers +43 660 1234567, 0316 123456 Starts with +, 00, 0 or a bracketed area code; 8 to 15 digits
Bank accounts (IBAN) AT61 1904 3002 3457 3201 Country code and the IBAN checksum (mod 97)
Card numbers 4111 1111 1111 1111 The Luhn checksum and the first digits of a real card scheme
US Social Security numbers 123-45-6789 style The format, with impossible ranges left out
UK National Insurance numbers 2 letters, 6 digits, 1 letter The format, with letters that are never used left out
VAT numbers ATU12345678, DE123456789 The format of each EU country’s number
Dates 03.10.2026, 2 November 2026 The format, and only real calendar dates; unticked by default
Mock of the Find personal data panel for the test invoice. Email addresses 2 of 2, one of them typed into a form field. Phone numbers 4 of 4, including an order number and a customer number that are not phone numbers. Bank accounts (IBAN) 1 of 1. Card numbers 1 of 1. VAT numbers 1 of 1. Dates 0 of 2. A button reads Mark 9 for redaction.Find personal dataFound by rules on this device. Nothing is removed until you save.Email addresses2 of 2Phone numbers4 of 4+43 660 12345670316 1234560042 1234 5678an order number: untick004312345678a customer number: untickBank accounts (IBAN)1 of 1Card numbers1 of 1VAT numbers1 of 1Dates0 of 2Mark 9 for redaction
The finder's list for our invented invoice: 9 hits ticked across 5 kinds, the 2 dates left unticked. The order number and the customer number are listed as phone numbers and need unticking.

What we found on a test invoice

We made an invented invoice with the usual suspects and ran the finder on it (11 October 2026):

  • Found and correct: the email address, the 2 phone numbers, the IBAN, the card number and the VAT number, and a second email typed into a form field. A hit in a field marks the whole field; saving emptied it, and the address was nowhere in the saved file.
  • Found, but not what they are: an order number “0042 1234 5678” and a customer number “004312345678”, both listed under Phone numbers, because they have the shape of 1. Untick them.
  • Found as dates: “03.10.2026” and “2 November 2026”, in the Dates group, unticked. Only real calendar dates count, so a reference like “01-23-4567” is not taken for one.
  • Rightly not found: an IBAN with 1 wrong digit. Its checksum fails, so it is not a valid account. It is still a person’s data if it is a typo of their real account, so a broken checksum is not a reason to relax.
  • Not found, as designed: the name “Jane Cooper”, the street address, and “the tenant in flat 4”. No rule finds those.

With all 9 ticked hits marked, the save took out 138 characters and 1 form field from 9 areas. On a 40-page bank statement with the IBAN and an email on every page and 30 dated lines a page, the finder listed the IBAN and the email 40 times each and the 1,200 dates under Dates, in about 0.1 seconds. Numbers that merely look like phone numbers are the false hits that remain, which is why the finder asks you to review instead of redacting on its own.

Where it does not look

The finder reads the text on the pages, the same text search finds (text on a hidden layer included), and the text typed into text fields and lists of a form. Personal data also sits in places it does not read:

  • Document properties. Author, title and the creating program. Our invoice still named its author after redaction. Clear them with Remove PDF metadata; what PDF metadata reveals lists what is there.
  • Attachments and bookmarks. A PDF can carry whole files, and a bookmark title is not page text. Redaction removes neither. Inspect lists attachments by name and bookmarks under Outline; hidden pages and attachments explains how to leave them behind.
  • Pictures. A photo of an ID, a signature, a scanned page without OCR. Text in a picture is pixels until OCR reads it.
  • Comments and notes that reviewers left on the pages.

Does it work on scanned pages?

Yes, after OCR. A scan has no text, so the finder finds nothing on it until OCR PDF adds a text layer. Then it reads what OCR recognised and boxes each hit at the height of the printed line, so a save paints the printed words out of the picture as well as taking out the invisible ones.

We measured it on 11 October 2026 on a scanned statement page (11-point type at 200 dpi). OCR read it at 94 % confidence; the finder listed both IBANs, the card number, the email and the mobile number, with boxes 10 to 13 points tall. After saving, we ran OCR on the redacted page again: none of those numbers came back.

Two reasons to look before you save. OCR can misread a digit, and a misread IBAN fails its checksum and is not listed. And a box sits on what OCR saw, so it can clip a letter of the word next to it. Redact a scanned PDF covers boxes drawn by hand on scans, and the OCR guide covers getting a good text layer first.

A manual sweep for what rules miss

Rules find shapes. You know the context. A 5-minute sweep catches the rest:

  1. In the editor, press Ctrl+F and search for every name you know is in the document, and for your organisation’s own customer prefixes.
  2. Search for words that sit next to personal data: “born”, “address”, “tenant”, “patient”, “account”, “salary”.
  3. Look at headers and footers on a few pages; a sender’s or recipient’s details often repeat there.
  4. Drop the file on Inspect for properties, attachments and comments, which search does not reach.

For a long file, PDF to text gives you all the text in 1 plain file to search in any editor.

Check the result

After saving, open the redacted file in any viewer and:

  1. Select all, copy and paste into a text editor. Search for each removed item.
  2. Search the PDF for the same items.
  3. Look at the pages where scans or pictures were redacted.

The editor writes the redacted file anew, so the removed text leaves no earlier copy behind in it. The privacy and protection guide covers what redaction removes and what a password protects.

The honest part: the finder shortens the search, not the decision

The GDPR defines personal data as “any information relating to an identified or identifiable natural person” (Article 4(1), checked on 11 October 2026). That reaches far beyond the patterns any finder can match: a job title in a small team, a flat number, a description that only fits 1 person. Whether a document may be shared, and with whom, depends on why you hold it and who receives it. We explain how the tools work; we do not give legal advice, and no tool replaces reading the document before it leaves.

What the finder does well is the tedious part: 40 IBANs in footers, every email in a long table, card numbers with valid checksums, all on your device, with nothing uploaded.

Questions

Does the finder find names?

No. Names and street addresses have no fixed shape a rule can check, so it does not try, and the panel says so. Search for the names you know with Ctrl+F and mark them by hand with the Redact tool.

Is anything removed when I click Mark?

No. Mark adds ordinary redaction boxes, which you can see, move or delete. Nothing leaves the file until you save, and 1 Ctrl+Z takes all the marks back.

Is my document uploaded to look for personal data?

No. The rules run in your browser tab, on your device. You can check: open the editor once so its engine is loaded, turn Wi-Fi off, and the finder still works.

Does it work on scanned PDFs?

Yes, once OCR PDF has added a text layer. The finder then reads what OCR recognised and boxes each hit at the height of the printed line, so saving paints the printed words out of the picture too. Look at the boxes before you save: a digit OCR misread is not found.

Does it make my PDF GDPR compliant?

No tool can. Deciding what counts as personal data in your context, and whether you may share it, is a judgement the finder cannot make. It shortens the search; the decision stays yours. This guide explains, it does not give legal advice.

The tools for this job