Find personal data in a PDF before you share it
By the getPDF team · Published 11 October 2026
The short answer
Open the PDF in Redact PDF and click Find personal data: it looks through every page on your device and lists email addresses, phone numbers, IBANs, card numbers, US Social Security, UK National Insurance and EU VAT numbers, and dates, in the page text and in filled form fields. Untick what may stay, click the Mark button, check the boxes and save; the marked text is taken out of the file. Names and street addresses are not found by rules, and the document’s properties, bookmarks and attachments are not searched, so check those yourself.
Run the finder
Redact PDFRemoves content, not just a black box. Free, runs on your device.- Drop the PDF on the page above, or open /redact-pdf. The editor opens with the Redact tool picked.
- Click Find personal data in the bar above the page. The panel shows “Looking through page 3 of 40…” while it reads.
- Read the list. Hits are grouped by kind, such as “Bank accounts (IBAN)”, with a count like “40 of 40” and the page of each hit.
- Untick anything that may stay. Dates start unticked, because most documents are full of harmless ones.
- Click Mark 9 for redaction (the number is what you left ticked). Each hit gets an ordinary redaction box on the page.
- Scroll through and look at the boxes. Add your own with the Redact tool (press R and drag) for anything the rules cannot find.
- Save with Ctrl+S. The status line says what was taken out, for example “Saved with 9 areas redacted: 138 characters, 1 form field taken out”.
Nothing is removed until you save, and every mark undoes. The file downloads with “-redacted” in its name; your original stays as it was.
What it finds, and how it checks
The finder uses rules, not guesses. Each kind has a shape, and the kinds with a checksum are checked, which keeps false hits down:
| Kind | Example it finds | How it is checked |
|---|---|---|
| Email addresses | jane.cooper@example.com | The shape name, at sign, domain |
| Phone numbers | +43 660 1234567, 0316 123456 | Starts with +, 00, 0 or a bracketed area code; 8 to 15 digits |
| Bank accounts (IBAN) | AT61 1904 3002 3457 3201 | Country code and the IBAN checksum (mod 97) |
| Card numbers | 4111 1111 1111 1111 | The Luhn checksum and the first digits of a real card scheme |
| US Social Security numbers | 123-45-6789 style | The format, with impossible ranges left out |
| UK National Insurance numbers | 2 letters, 6 digits, 1 letter | The format, with letters that are never used left out |
| VAT numbers | ATU12345678, DE123456789 | The format of each EU country’s number |
| Dates | 03.10.2026, 2 November 2026 | The format, and only real calendar dates; unticked by default |
What we found on a test invoice
We made an invented invoice with the usual suspects and ran the finder on it (11 October 2026):
- Found and correct: the email address, the 2 phone numbers, the IBAN, the card number and the VAT number, and a second email typed into a form field. A hit in a field marks the whole field; saving emptied it, and the address was nowhere in the saved file.
- Found, but not what they are: an order number “0042 1234 5678” and a customer number “004312345678”, both listed under Phone numbers, because they have the shape of 1. Untick them.
- Found as dates: “03.10.2026” and “2 November 2026”, in the Dates group, unticked. Only real calendar dates count, so a reference like “01-23-4567” is not taken for one.
- Rightly not found: an IBAN with 1 wrong digit. Its checksum fails, so it is not a valid account. It is still a person’s data if it is a typo of their real account, so a broken checksum is not a reason to relax.
- Not found, as designed: the name “Jane Cooper”, the street address, and “the tenant in flat 4”. No rule finds those.
With all 9 ticked hits marked, the save took out 138 characters and 1 form field from 9 areas. On a 40-page bank statement with the IBAN and an email on every page and 30 dated lines a page, the finder listed the IBAN and the email 40 times each and the 1,200 dates under Dates, in about 0.1 seconds. Numbers that merely look like phone numbers are the false hits that remain, which is why the finder asks you to review instead of redacting on its own.
Where it does not look
The finder reads the text on the pages, the same text search finds (text on a hidden layer included), and the text typed into text fields and lists of a form. Personal data also sits in places it does not read:
- Document properties. Author, title and the creating program. Our invoice still named its author after redaction. Clear them with Remove PDF metadata; what PDF metadata reveals lists what is there.
- Attachments and bookmarks. A PDF can carry whole files, and a bookmark title is not page text. Redaction removes neither. Inspect lists attachments by name and bookmarks under Outline; hidden pages and attachments explains how to leave them behind.
- Pictures. A photo of an ID, a signature, a scanned page without OCR. Text in a picture is pixels until OCR reads it.
- Comments and notes that reviewers left on the pages.
Does it work on scanned pages?
Yes, after OCR. A scan has no text, so the finder finds nothing on it until OCR PDF adds a text layer. Then it reads what OCR recognised and boxes each hit at the height of the printed line, so a save paints the printed words out of the picture as well as taking out the invisible ones.
We measured it on 11 October 2026 on a scanned statement page (11-point type at 200 dpi). OCR read it at 94 % confidence; the finder listed both IBANs, the card number, the email and the mobile number, with boxes 10 to 13 points tall. After saving, we ran OCR on the redacted page again: none of those numbers came back.
Two reasons to look before you save. OCR can misread a digit, and a misread IBAN fails its checksum and is not listed. And a box sits on what OCR saw, so it can clip a letter of the word next to it. Redact a scanned PDF covers boxes drawn by hand on scans, and the OCR guide covers getting a good text layer first.
A manual sweep for what rules miss
Rules find shapes. You know the context. A 5-minute sweep catches the rest:
- In the editor, press Ctrl+F and search for every name you know is in the document, and for your organisation’s own customer prefixes.
- Search for words that sit next to personal data: “born”, “address”, “tenant”, “patient”, “account”, “salary”.
- Look at headers and footers on a few pages; a sender’s or recipient’s details often repeat there.
- Drop the file on Inspect for properties, attachments and comments, which search does not reach.
For a long file, PDF to text gives you all the text in 1 plain file to search in any editor.
Check the result
After saving, open the redacted file in any viewer and:
- Select all, copy and paste into a text editor. Search for each removed item.
- Search the PDF for the same items.
- Look at the pages where scans or pictures were redacted.
The editor writes the redacted file anew, so the removed text leaves no earlier copy behind in it. The privacy and protection guide covers what redaction removes and what a password protects.
The honest part: the finder shortens the search, not the decision
The GDPR defines personal data as “any information relating to an identified or identifiable natural person” (Article 4(1), checked on 11 October 2026). That reaches far beyond the patterns any finder can match: a job title in a small team, a flat number, a description that only fits 1 person. Whether a document may be shared, and with whom, depends on why you hold it and who receives it. We explain how the tools work; we do not give legal advice, and no tool replaces reading the document before it leaves.
What the finder does well is the tedious part: 40 IBANs in footers, every email in a long table, card numbers with valid checksums, all on your device, with nothing uploaded.
Questions
Does the finder find names?
No. Names and street addresses have no fixed shape a rule can check, so it does not try, and the panel says so. Search for the names you know with Ctrl+F and mark them by hand with the Redact tool.
Is anything removed when I click Mark?
No. Mark adds ordinary redaction boxes, which you can see, move or delete. Nothing leaves the file until you save, and 1 Ctrl+Z takes all the marks back.
Is my document uploaded to look for personal data?
No. The rules run in your browser tab, on your device. You can check: open the editor once so its engine is loaded, turn Wi-Fi off, and the finder still works.
Does it work on scanned PDFs?
Yes, once OCR PDF has added a text layer. The finder then reads what OCR recognised and boxes each hit at the height of the printed line, so saving paints the printed words out of the picture too. Look at the boxes before you save: a digit OCR misread is not found.
Does it make my PDF GDPR compliant?
No tool can. Deciding what counts as personal data in your context, and whether you may share it, is a judgement the finder cannot make. It shortens the search; the decision stays yours. This guide explains, it does not give legal advice.
The tools for this job
Keep reading
- Redact a PDF properly: remove the text, never just cover itA black rectangle hides nothing: the text underneath still copies out.
- Redact a bank statement before you share it, and keep what they needLandlords and visa offices need your name, dates and balance, not your card number and every purchase.
- GDPR and online PDF tools: what is allowed at work, in plain wordsGDPR and online PDF tools: uploading a client PDF to a converter is processing personal data.
- Clean a PDF before publishing it: the 7-step pass for hidden leftoversMetadata, comments, bookmarks, attachments, old form data, earlier versions: the 7-step pass that catches what a visual check misses before a PDF goes public.
- PDF privacy and protection: the complete guideWhat a PDF password really protects, how true redaction works, what metadata leaks, and how to send files safely.
- Hidden pages, layers and attachments in a PDF: what can hide, and how to find itA PDF can carry attached files, invisible layers, earlier versions and pages no viewer shows.