Is the PDF a scan or real text? Check it before you convert
By the getPDF team · Published 11 October 2026
The short answer
Drop the PDF on Inspect below. “Pages with text” against the page count, and “Scanned pages (image, no text)”, tell you in a few seconds whether the file is real text, a scan or a mix. Real text converts straight to Word, Excel, text or Markdown; scanned pages need OCR PDF first; “Copy allowed: no” means the owner restricted copying. Then run PDF to Text on the file and read 1 paragraph, because some problems only show in the text itself.
Try it here, nothing is uploaded
The check without any tool
- Select a word. Drag across a line in your PDF viewer. If single words highlight, there is text. If nothing highlights, or 1 box covers the whole page, there is probably none.
- Search for a word you can see. Ctrl+F (Cmd+F on a Mac). No match for a word that is plainly on the page means no text.
- Zoom in to 400 %. Real text stays crisp at any zoom, because it is drawn from a font. Scanned letters go soft and blocky, and you often see paper grain or a slight tilt.
That settles most files. It misses 2 cases: a long file where only some pages are scanned, and text that selects fine but is broken inside. Inspect catches the first; reading the text catches the second.
What Inspect tells you, and what it means for converting
Inspect reads the file and changes nothing. These are the rows that matter before a conversion, with the readings from our test files:
| Row in Inspect | Typed invoice | Scanned lease | Contract with a scanned page |
|---|---|---|---|
| Pages | 3 | 1 | 4 |
| Pages with text | 3 of 3 checked | 0 of 1 checked | 3 of 4 checked |
| Scanned pages (image, no text) | 0 | 1 | 1 |
| Fonts | Helvetica, Helvetica-Bold | none (no text) | Helvetica, Helvetica-Bold |
| Images, share of the file | 0, 0 % | 1, 97 % | 1, 95 % |
| Copy allowed | yes | yes | yes |
| What to do | convert | OCR first | OCR, then convert |
How to read each row:
- Pages with text equal to the page count: every page has real text and will convert.
- Scanned pages above 0: those pages are pictures. A converter gives you nothing for them (PDF to Text) or the picture itself (PDF to Word). Run OCR PDF first; it reads only the pages without text and leaves the others alone.
- Fonts “not embedded” matters for how the PDF looks on other devices, not for conversion: the text is still there.
- Images taking 90 % or more of the file, with few or no text pages, is the signature of a scan.
- Copy allowed: no, with “Encrypted: yes”: the owner restricted copying. That is a permission question first, a technical one second; the guide to text that will not copy covers it.
Inspect checks the first 200 pages for text, fonts and pictures and counts the rest. On our 500-page file it took 204 ms and said “Scanned the first 200 of 500 pages for fonts and images”.
The mixed file, the common surprise
A typed contract with 1 page that was printed, signed and scanned back in. Everything converts except that page, and a conversion of 40 pages can hide it well.
Inspect shows it as a count: on our 4-page test file, “Pages with text: 3 of 4 checked” and 1 scanned page. It does not say which. PDF to text does: its result warns “has no text on page 4: probably scanned. Run OCR PDF first to make the text readable, then convert it.” PDF to Word puts such a page into the document as its picture and gives the same kind of warning.
Pick the route
- Real text, copying allowed: convert. PDF to Word for prose to edit, PDF to Excel for tables, PDF to text or PDF to Markdown for notes, search and AI tools. The conversion guide helps pick.
- Scanned pages: OCR PDF first, then convert. The scan guide covers languages and confidence.
- Copy allowed: no: Unlock PDF, for a file you have the right to use.
- Text, but it reads wrong: see the next section.
The test Inspect cannot do: read the text
Inspect predicts; it does not guarantee. Two problems only show in the text itself:
- A font without its character map. We built a file whose embedded font had lost its map from shapes back to letters. Inspect showed the font as embedded and the page as having text, all true and all normal. The text came out as
ăµGļRIILFHWKHOHDVHVWDUWVRQ0DUFKfor “Łódź office: the lease starts on 1 March.” - Text stored out of reading order. A 2-column page whose program drew it line by line across both columns comes out with the columns mixed, in every converter. Inspect has no row for order.
So before a big job, run PDF to text on the file and read 1 paragraph from the most complicated page. It is quick: a 500-page file took 173 ms in our test. If it reads right, convert with confidence. If it reads wrong, the guide to jumbled PDF text says which fix applies, and for a 1-page letter, retyping may simply be faster.
Everything on this page runs in your browser tab: the file is read on your device and not uploaded.
Questions
How can I tell if a PDF is scanned without any tool?
Try to select a word, and search for a word you can see. If nothing selects and search finds nothing, it is a scan. Zoom in to 400 %: real text stays sharp, scanned letters go soft and blocky.
What does Inspect count as a scanned page?
A page with no text that holds at least 1 picture. A page that is only a photo or a diagram counts too, so read it as no text on this page. A page with no text and no picture, such as letters converted to outlines, counts as neither.
My PDF has text. Will it convert well?
Probably, but not certainly. A file can have text that comes out as gibberish (a font without its character map) or in a mixed order (2 columns stored line by line). Inspect cannot see either. Run PDF to Text on it and read a paragraph: that takes seconds.
Does Inspect check every page?
It scans the first 200 pages for text, fonts and pictures and counts the rest. On our 500-page test file it said so: Scanned the first 200 of 500 pages. Running PDF to Text on the whole file checks every page.
The tools for this job
Keep reading
- Cannot select or copy text in a PDF: the 3 causes and their fixesText in a PDF will not select or copy for 1 of 3 reasons: a scan, a copy restriction, or a font that hides the letters.
- Text copied from a PDF comes out jumbled: why, and what fixes itCopy and paste from a PDF comes out jumbled: wrong order, gibberish or glued words.
- Convert a PDF to Word you can edit, not a page of stuck text boxesConvert a PDF to a Word file you can really edit: flowing paragraphs, Word headings, lists and tables, not stacked text boxes.
- Convert many PDFs to text at once, and catch the scans that come out emptyConvert multiple PDFs to text in one batch in your browser: a whole folder in, 1 .txt per PDF out, and how to spot the scanned files that come out empty.
- Make a scanned PDF searchable, without changing how it looksOCR a scanned PDF on your device: free, no account, nothing is uploaded.
- Get the text out of a PDF for translation, clean enough to quote and translateExtract text from a PDF for translation: a Word file with whole paragraphs, no page numbers or headers in the flow, and an honest word count for the quote.
- PDF to Word, Excel and text: what converts, what breaks, and how to pickConvert PDF to Word, Excel, plain text, Markdown or CSV free in your browser.