Clean a PDF before publishing it: the 7-step pass for hidden leftovers
By the getPDF team · Published 11 October 2026
The short answer
Before a PDF goes on a website or outside your organisation, list everything in it with Inspect, redact what must go for real, flatten or delete comments and form fields, write a fresh copy that drops attachments, bookmarks and earlier versions, remove the metadata, rename the file, and check the final copy cold. Seven steps, a few minutes, all on your device. A visual check misses most of this, because none of it is on the pages.
Why looking at the pages is not enough
A PDF can carry far more than its pages show. In our tests, a 4-page file with an author name, a title, 4 bookmarks, an XMP packet and an attached spreadsheet looked like 4 ordinary pages in every viewer. And a file edited and saved the way many editors save (by appending the change) still held the old invoice total in its bytes, though no viewer showed it.
Published-document incidents almost always trace back to 1 skipped step. In the 2011 UK submarine report, the redacted passages were still text in the file, and a search engine made them readable to anyone. Each step below exists for a reason like that.
The 7 steps at a glance
| Step | What it catches | Tool | Done when |
|---|---|---|---|
| 1. List everything | You cannot remove what you have not seen | Inspect | You have read every panel |
| 2. Redact for real | Names, numbers, pictures that must not be public | Redact PDF | Copy and paste returns none of it |
| 3. Comments and forms | Comment authors, old field values | Flatten PDF, or delete in the editor | Annotations 0, form none |
| 4. A fresh copy | Attachments, bookmarks, earlier versions, stray pages | Extract pages, Delete pages | Outline and Attachments say none |
| 5. Metadata | Author, programs, dates, XMP | Remove PDF metadata | The card says what it removed |
| 6. File name | A name that says more than the title | Your file manager | The name is neutral |
| 7. Cold check | Anything the steps above missed | Inspect, PDF to Text | A stranger’s read finds nothing |
Step 1: list everything with Inspect
Drop the file on Inspect. Nothing is changed and nothing is uploaded. Read:
- Metadata: Title, Author, Creator, Producer, dates.
- XMP metadata: the second copy of the title, author, tools and dates, which a properties dialog often does not show.
- Outline: every bookmark title. Bookmarks are often written early (“Section 4, figures from client X”) and never revisited.
- Security panel: Annotations (comments, highlights, links), Form, Digital signatures, Scripts, and Attachments with their file names.
- Pages: the count. A 12-page file where you expected 10 has 2 pages to find.
Write down what must go. Inspect lists attachments by name but does not open them.
Step 2: redact what must go, for real
If anything on the pages must not be public, open the file at /redact-pdf, drag a box over each item (or click Find personal data for emails, phone numbers, IBANs and card numbers), and save. The characters, picture pixels and drawings under each box leave the file, and the editor writes it anew. A black rectangle from a drawing tool does not do this; Redact a PDF properly shows the difference.
Step 3: comments and form fields
Comments carry their authors’ names, and a form can still hold the values the last person typed.
- To keep comments visible but fixed, run Flatten PDF: it draws form fields and annotations into the page content and removes the fields. It draws what is in the fields, so clear any old values first. A note keeps only its icon; its text leaves the file.
- To remove a comment or a highlight, open the file in the editor, select it and delete it.
- To remove an old value, clear the field in the editor before you flatten, or redact it.
Flattening also ends any digital signature, as every change does.
Step 4: a fresh copy without attachments, bookmarks and old versions
Run Extract pages on the file with Pages set to “1-” (every page from 1 to the end). It copies the pages into a new file and nothing else. We measured what that leaves out:
- the attached spreadsheet: gone;
- the 4 bookmarks: gone;
- an earlier version of an edited page that an appending save had kept in the file: gone (the old total was in the input’s bytes and not in the output’s).
If the file has pages that should not be published (a cover sheet with a distribution list, a blank page with notes), drop them with Delete pages, or list only the pages you want in Extract pages.
Skip this step for a file whose attachments are the point, such as an e-invoice that must carry its XML.
Step 5: remove the metadata
Drop the fresh copy on Remove PDF metadata and click Remove metadata. It removes every document property, the XMP packet, per-page metadata and private application data, and replaces the file identifier. Do this after step 4, because the fresh copy names getPDF as its producer, and this step removes that too.
It also removes leftovers nothing points at, such as an earlier page version from an appending save, and the card says how many. Still, Remove PDF metadata on its own is not a full clean: in our test it left bookmarks and attachments exactly where they were. That is why step 4 comes first.
Step 6: rename the file
The file name travels with the PDF and shows in every download bar: “offer-v3-for-graz-client-DO-NOT-SEND.pdf”. Give it a name you would print on the first page.
Step 7: open the final copy cold and check it
- Inspect it again. Metadata and XMP metadata none, Outline none (or titles you chose), Annotations and Attachments 0, Form none if you flattened, pages as expected.
- Search the text. Run PDF to Text and search for every name, number and word you removed, and for your organisation’s name.
- Read it as an outsider. Open it in a different viewer from the one you work in, read the first and last pages, and look at the sidebar.
The honest part
This pass cleans the file. Whether the content itself should be public is a different review, and no tool does it for you: a published table can identify people even with every name removed, and a confidential plan stays confidential only if it is not posted.
The tools have edges too. Remove PDF metadata does not reach metadata inside embedded pictures (a photo’s camera details). Extract pages does not keep the bookmarks you might want for a long public report; add them back in the source document instead. And flattening a filled form publishes whatever was in the fields, so read it before you run it. If a step’s result surprises you, stop and inspect again before you publish.
Questions
What hidden data can a PDF contain?
Document properties and an XMP packet (author, programs, dates), comments with their authors' names, bookmarks, attached files, form fields with old values, and, after some editors' saves, earlier versions of the pages kept inside the file.
Does Remove PDF metadata clean everything?
No. It removes the document properties, the XMP packet, private application data and leftovers nothing points at, such as earlier versions of edited pages. Comments, bookmarks and attachments stay; this pass covers those with other tools.
What is an earlier version inside a PDF?
Some editors save changes by appending them to the end of the file. The old page content stays in the file, unused, where a tool that reads every byte can find it. Writing the file anew drops it, and so does Remove PDF metadata.
Is my file uploaded at any step?
No. Inspect, Redact PDF, Flatten PDF, Extract pages and Remove PDF metadata all run in your browser tab; nothing is uploaded.
The tools for this job
Keep reading
- What the metadata in your PDFs reveals, and how to remove itAuthor names, employers, software versions, edit dates: the metadata a typical PDF carries, how to read yours in 10 seconds, and how to remove it all.
- Hidden pages, layers and attachments in a PDF: what can hide, and how to find itA PDF can carry attached files, invisible layers, earlier versions and pages no viewer shows.
- Redact a PDF properly: remove the text, never just cover itA black rectangle hides nothing: the text underneath still copies out.
- What flattening a PDF does, and what it never protectedFlattening turns form fields and annotations into fixed page content.
- PDF redaction failures that made the news, and how each one happenedThe 2005 Calipari report, the 2011 UK submarine papers, the 2019 Manafort filing: how each redaction failed, and the 30-second check that catches it.
- PDF privacy and protection: the complete guideWhat a PDF password really protects, how true redaction works, what metadata leaks, and how to send files safely.