Skip to content

What the metadata in your PDFs reveals, and how to remove it

By the getPDF team · Published 11 October 2026

The short answer

A PDF usually carries a few lines describing itself: who wrote it (often your full name, taken from your Office user name), the title, the program and version that made it, and when it was created and last changed. Drop a file on Inspect to see these, and the XMP copy most quick cleaners miss, in 10 seconds; Remove PDF metadata clears both. Both run on your device. Metadata is only 1 place a PDF can reveal things; the content is another.

Inspect PDFFonts, images, metadata, why it is big. Free, runs on your device.

The standard fields, and what each gives away

Every PDF can have a document information dictionary, the list your viewer shows under File, Properties. These are its fields, and what they typically say about you:

Field What it holds What it can reveal
Author A person’s name Your full name, or your login name at work, filled in by the program, not by you
Title The document’s title An old working title (“Offer v3 FINAL for Graz”) that says more than the file name
Subject, Keywords Free text Internal project names, client names, tags
Creator The program the document was written in Your toolchain and often its version; an internal or licensed tool names your employer
Producer The program that wrote the PDF The PDF library or printer driver, with a version
CreationDate When the file was made A timeline: that a “fresh” offer was made 8 months ago
ModDate When it was last changed That a document was edited after it was signed or sent

None of this is visible on the pages. All of it is visible to anyone who opens File, Properties, and to every search engine that indexes a published PDF.

A worked example

We ran our metadata test file, an invented “Quarterly report”, through Inspect. The Metadata panel listed 8 fields:

Field Value
Title Quarterly report
Author Jane Cooper
Subject Q3 figures
Keywords finance q3 (and 1 more tag)
Creator Studio Lumen Sheets
Producer Studio Lumen PDF 4.1
CreationDate 2026-07-01 09:00 UTC
ModDate 2026-09-15 14:30 UTC

Inside the file the dates are stored as D:20260701090000Z (D, then year, month, day, hour, minute, second, and Z for UTC); Inspect turns them into dates you can read. Its XMP metadata panel listed the title and Jane Cooper a second time, from the file’s other store. Read together, the 8 lines say who wrote it, in which program, for which project, and that it sat for 11 weeks between creation and the last edit.

XMP: the second store most quick cleaners miss

Since PDF 1.4 a file can also carry an XMP packet: a block of XML stored separately from the information dictionary, holding the same facts and often more (some editors add a history of saves and the tools used). The PDF 2.0 standard, ISO 32000-2, deprecates most of the old dictionary in favour of XMP, so newer files lean on it more.

The trap: a tool or a properties dialog that clears the Author field in the dictionary can leave the XMP copy untouched, with your name still in it. Our test file has exactly that: Jane Cooper as Author in the dictionary and again as creator in the XMP packet.

Inspect shows both stores: the dictionary under Metadata, and the XMP packet’s title, author, creator tool, producer and dates under XMP metadata. Remove PDF metadata clears both in 1 pass. On the test file its result card said: “Removed 8 document properties”, “Removed 1 XMP metadata packet”, “Replaced the file identifier”. Afterwards both of Inspect’s metadata panels said none, and a search of the cleaned file’s raw bytes found no trace of “Jane”.

Remove metadataAuthor, dates, XMP, hidden data. Free, runs on your device.

It also removes per-page XMP packets and PieceInfo, where some applications keep private data, and it replaces the file identifier, which can be derived from the original’s contents and dates. It also removes leftovers nothing points at any more, such as earlier versions of pages edited with an appending save, and the card says how many. You can keep the title if you want readers to see it in their tab; everything else goes.

Photos inside PDFs keep their own metadata

A JPEG from a phone or camera carries EXIF data: the device model, the time, and often GPS coordinates of where it was taken. When a program puts that photo into a PDF without re-encoding it, the EXIF block often goes in too.

Our own JPG to PDF copies the picture data unchanged, so the quality stays, but leaves out the camera’s metadata (location, device, time) from each JPEG, and the result says how many photos it cleaned. A PDF made elsewhere may still carry it, and Extract images hands a JPEG back byte for byte, so whatever EXIF was inside comes out with it.

Removing PDF metadata does not touch pictures already inside a PDF; it describes the file, not the pictures in it. If a photo’s location matters and the PDF comes from another program, strip it from the photo before making the PDF. On Windows: right-click the photo, Properties, Details, “Remove Properties and Personal Information”. Most phones also let you turn off location when sharing a photo.

What to do before a PDF leaves your hands

  1. Inspect it. Drop it on Inspect and read the Metadata and XMP metadata panels. Ten seconds.
  2. Clean it with Remove PDF metadata if any line says more than you want to share.
  3. Inspect the clean copy to confirm both panels say “none”.
  4. Fix the source for next time. In Word, File, Info, Check for Issues, Inspect Document removes document properties and personal information from the Word file before you export it, as Microsoft’s support pages describe (checked on 11 October 2026).

For a full pass before publishing, comments, attachments and form data included, see the privacy and protection guide.

The honest part

Metadata removal matters for files you publish or send to people outside your organisation. It does nothing about the content: a name in a header, a signature line, a phone number in the footer or an address in the body stays exactly where it is, and so do comments and attachments. Inspect lists a file’s attachments by name, but not what is in them; Hidden pages and attachments covers that layer. And a clean file can still identify you through what it says; Anonymise a PDF deals with that, and removing content for good is a job for redaction, not for a metadata tool.

Questions

What metadata does a PDF contain?

Usually a short list of document properties (title, author, subject, keywords, the program that made it, the program that wrote the PDF, a creation date and a modification date), often an XMP packet repeating them and more, sometimes private application data and a file identifier.

Where does my name in the Author field come from?

Mostly from the user name in Word, Excel or your operating system, which the program writes into Author when it saves or exports. Most people never typed it.

Does removing metadata change the pages?

No. Text, images and layout stay exactly as they were. Only the entries that describe the file are removed, and the file gets a new identifier.

Can photos inside a PDF reveal where they were taken?

Yes. A JPEG photo can carry EXIF data, including GPS coordinates, and many programs embed photos byte for byte, EXIF and all. Our JPG to PDF leaves that block out and says how many photos it cleaned. Removing PDF metadata does not touch pictures already inside a PDF.

Is the file uploaded to read or clean it?

No. Inspect and Remove metadata run in your browser tab; nothing is uploaded. Use either once, turn Wi-Fi off, and both still work.

The tools for this job