Skip to content

Is the PDF a scan or real text? Check it before you convert

By the getPDF team · Published 11 October 2026

The short answer

Drop the PDF on Inspect below. “Pages with text” against the page count, and “Scanned pages (image, no text)”, tell you in a few seconds whether the file is real text, a scan or a mix. Real text converts straight to Word, Excel, text or Markdown; scanned pages need OCR PDF first; “Copy allowed: no” means the owner restricted copying. Then run PDF to Text on the file and read 1 paragraph, because some problems only show in the text itself.

Try it here, nothing is uploaded

PDF · any size

The check without any tool

  1. Select a word. Drag across a line in your PDF viewer. If single words highlight, there is text. If nothing highlights, or 1 box covers the whole page, there is probably none.
  2. Search for a word you can see. Ctrl+F (Cmd+F on a Mac). No match for a word that is plainly on the page means no text.
  3. Zoom in to 400 %. Real text stays crisp at any zoom, because it is drawn from a font. Scanned letters go soft and blocky, and you often see paper grain or a slight tilt.

That settles most files. It misses 2 cases: a long file where only some pages are scanned, and text that selects fine but is broken inside. Inspect catches the first; reading the text catches the second.

What Inspect tells you, and what it means for converting

Inspect reads the file and changes nothing. These are the rows that matter before a conversion, with the readings from our test files:

Row in Inspect Typed invoice Scanned lease Contract with a scanned page
Pages 3 1 4
Pages with text 3 of 3 checked 0 of 1 checked 3 of 4 checked
Scanned pages (image, no text) 0 1 1
Fonts Helvetica, Helvetica-Bold none (no text) Helvetica, Helvetica-Bold
Images, share of the file 0, 0 % 1, 97 % 1, 95 %
Copy allowed yes yes yes
What to do convert OCR first OCR, then convert

How to read each row:

  • Pages with text equal to the page count: every page has real text and will convert.
  • Scanned pages above 0: those pages are pictures. A converter gives you nothing for them (PDF to Text) or the picture itself (PDF to Word). Run OCR PDF first; it reads only the pages without text and leaves the others alone.
  • Fonts “not embedded” matters for how the PDF looks on other devices, not for conversion: the text is still there.
  • Images taking 90 % or more of the file, with few or no text pages, is the signature of a scan.
  • Copy allowed: no, with “Encrypted: yes”: the owner restricted copying. That is a permission question first, a technical one second; the guide to text that will not copy covers it.

Inspect checks the first 200 pages for text, fonts and pictures and counts the rest. On our 500-page file it took 204 ms and said “Scanned the first 200 of 500 pages for fonts and images”.

The mixed file, the common surprise

A typed contract with 1 page that was printed, signed and scanned back in. Everything converts except that page, and a conversion of 40 pages can hide it well.

Inspect shows it as a count: on our 4-page test file, “Pages with text: 3 of 4 checked” and 1 scanned page. It does not say which. PDF to text does: its result warns “has no text on page 4: probably scanned. Run OCR PDF first to make the text readable, then convert it.” PDF to Word puts such a page into the document as its picture and gives the same kind of warning.

Pick the route

Decision tree. Start: does text select? If not, or Inspect shows scanned pages: OCR first. If copy allowed is no: unlock a file you have the right to use. If the text reads correctly in PDF to Text: convert. If it is gibberish or mixed: OCR with read again, or copy by column, or retype.Inspect the PDFScanned pages above 0?yesOCR firstOCR PDF, then convertnoCopy allowed: no?yesUnlock firstif you have the right tonoDoes PDF to Text read right?yesnoConvertWord, Excel, text, MarkdownGibberish: OCR, read againMixed order: copy by column, or retype
From a 1-minute check to the right tool. Inspect answers the first 2 questions; 1 run of PDF to Text answers the third.
  1. Real text, copying allowed: convert. PDF to Word for prose to edit, PDF to Excel for tables, PDF to text or PDF to Markdown for notes, search and AI tools. The conversion guide helps pick.
  2. Scanned pages: OCR PDF first, then convert. The scan guide covers languages and confidence.
  3. Copy allowed: no: Unlock PDF, for a file you have the right to use.
  4. Text, but it reads wrong: see the next section.

The test Inspect cannot do: read the text

Inspect predicts; it does not guarantee. Two problems only show in the text itself:

  • A font without its character map. We built a file whose embedded font had lost its map from shapes back to letters. Inspect showed the font as embedded and the page as having text, all true and all normal. The text came out as ăµGļRIILFHWKHOHDVHVWDUWVRQ0DUFK for “Łódź office: the lease starts on 1 March.”
  • Text stored out of reading order. A 2-column page whose program drew it line by line across both columns comes out with the columns mixed, in every converter. Inspect has no row for order.

So before a big job, run PDF to text on the file and read 1 paragraph from the most complicated page. It is quick: a 500-page file took 173 ms in our test. If it reads right, convert with confidence. If it reads wrong, the guide to jumbled PDF text says which fix applies, and for a 1-page letter, retyping may simply be faster.

Everything on this page runs in your browser tab: the file is read on your device and not uploaded.

Questions

How can I tell if a PDF is scanned without any tool?

Try to select a word, and search for a word you can see. If nothing selects and search finds nothing, it is a scan. Zoom in to 400 %: real text stays sharp, scanned letters go soft and blocky.

What does Inspect count as a scanned page?

A page with no text that holds at least 1 picture. A page that is only a photo or a diagram counts too, so read it as no text on this page. A page with no text and no picture, such as letters converted to outlines, counts as neither.

My PDF has text. Will it convert well?

Probably, but not certainly. A file can have text that comes out as gibberish (a font without its character map) or in a mixed order (2 columns stored line by line). Inspect cannot see either. Run PDF to Text on it and read a paragraph: that takes seconds.

Does Inspect check every page?

It scans the first 200 pages for text, fonts and pictures and counts the rest. On our 500-page test file it said so: Scanned the first 200 of 500 pages. Running PDF to Text on the whole file checks every page.

The tools for this job