Skip to content

Convert many PDFs to text at once, and catch the scans that come out empty

By the getPDF team · Published 11 October 2026

The short answer

Open PDF to Text, drop a folder of PDFs on it (or press “or a whole folder”), and press Get the text. Each PDF becomes a .txt file with the same name, done one after another, so one damaged or locked file does not stop the rest. Download them all as 1 zip, or in Chrome and Edge save them straight into a folder. Then look at the sizes: a text file of a few bytes came from a scan, which needs OCR first.

Try it here, nothing is uploaded

PDF · any size · many at once

Convert a folder of PDFs to text, step by step

  1. Drop the folder on the tool above, or press or a whole folder under the drop zone and pick it. Subfolders are included. Files that are not PDFs are left out, and the page says how many (“Left out 3 files this tool does not take”).
  2. Check the list. Each file shows its page count and size; a file that needs a password to open shows “locked”.
  3. Pick what goes between pages under Between pages: Empty line (the default), Page 2 (a line naming the page that follows, useful when you need to cite pages later), or Page break (a form feed character that printers and some text tools split on).
  4. Press Get the text. The progress line reads like “Reading the text · file 12 of 200 · page 2 of 3”.
  5. Read the result. It says “200 files are ready”, lists every .txt with its size, and names each file’s character count, for example “invoice.pdf: Took 172 characters of text from 3 pages, in the order a PDF reader finds it, as UTF-8”. Files that could not be done are listed under “1 file was not done; the others were:” with the reason.
  6. Save. Download all (200) as zip gives you getpdf-files.zip. In Chrome and Edge, Save to a folder writes each .txt straight into a folder you pick, with no zip in between. Nothing already in the folder is replaced: if report.txt is there, the new one is saved as report-2.txt. An empty folder is still the tidiest choice.

The text is UTF-8, so accents, Polish, Czech, Greek and Cyrillic come through as they are: “Łódź” stays “Łódź”.

How many files can one batch take?

The batch runs each PDF as its own job, one after another, and only 1 file is open in the engine at a time. That is why the count barely matters. We measured the same code the page runs, on our test laptop:

  • 200 invoices of 3 pages each (0.8 MB of PDFs) became 1.9 million characters of text in 6.7 seconds, about 33 ms a file.
  • 1 file of 500 pages became 21,891 characters of text in 173 ms.

In the browser tab, expect a little more time per file than that. What does not scale is a single gigantic file. The engine has to hold the whole PDF in the browser’s memory while it reads it; if a file does not fit, that file is listed as “too big for this browser to open”, with what the device can hold, and the batch carries on with the next one. Text itself is small, so hundreds of results wait in the tab without trouble until you save them.

The silent failure: scans that come out empty

A scanned PDF converts “successfully” into nothing, because its pages are pictures of text, not text. These are real results from our test files:

PDF Pages Text file What it tells you
3-page typed invoice 3 172 B real text, fine
500-page text file 500 21,891 B real text, fine
scanned lease 1 1 B (a single line break) a scan: no text at all
3-page contract with a scanned page added 4 174 B page 4 is a scan
3-page file with photos 3 70 B page 2 is only a photo

The tool does not hide this. Each such file gets a warning in the result, worded like “scan-lease.pdf has no text at all: probably scanned. Run OCR PDF first to make the text readable, then convert it”, or for a mixed file “has no text on page 4”. A page that is only a photo or a diagram gets the same note, so read it as “nothing to read here” rather than proof of a scan.

To find them in a big batch, look down the size column in the result list: anything of a few bytes, or far smaller than files of a similar page count, is your list for OCR.

Send the scans through OCR, then convert again

  1. Collect the PDFs the warnings named.
  2. Open OCR PDF, drop them all, pick the language they are written in, and press Make it searchable. It adds an invisible text layer and leaves the page pictures as they were.
  3. Drop the OCR results on PDF to Text again.

OCR is a guess, and the tool says how sure it is for each page. On a scanned bank statement it averaged 94 % confidence in our tests. The guide to making scans searchable covers languages, confidence and handwriting.

Text or Markdown for an AI pipeline or a search index

PDF to Text keeps each printed line as a line and keeps everything on the page, including the running header and the page number on every page. That is right for search indexes and for diffing, less right for anything that reads meaning.

PDF to Markdown takes the same batch and rebuilds the structure: headings from larger font sizes, paragraphs from line spacing, lists from bullets and numbers, tables as Markdown tables, and words broken by a hyphen at a line end joined again. Running lines (headers, footers and page numbers that repeat on most pages) are left out by default. On our 4-page test lease, the plain text counted 1,132 words and the Markdown 1,092: the 40 words of difference are the header and “Page 1 of 4” lines that every chunk of a pipeline would otherwise carry. If you feed the files to a language model or a notes app, Markdown is the better batch.

Spot-check 1 hard file before trusting the batch

The tools take the text in the order the file stores it, which is the order the creating program drew it. For letters, reports and most exports that is the reading order. For some 2-column layouts it is not: on a test page whose program drew the 2 columns line by line across the page, the text came out with both columns mixed, line by line, and Markdown and Word followed the same order. In the .txt, tables come out as lines of text; PDF to Markdown writes them as Markdown tables.

So before you rely on 2,000 files, open the text of the most complicated one (a newsletter, a 2-column paper, a form) next to its PDF. If it reads right, the batch is probably fine; if it does not, the guide to jumbled PDF text explains what you are seeing and what can and cannot fix it.

Two more limits. A PDF with a password to open is skipped in a batch (“It has a password. Run it on its own to type the password.”), and the folder structure is not kept: every .txt lands in 1 zip or 1 folder.

Questions

How many PDFs can I convert at once?

There is no count limit. Files are done one after another, so only 1 is open at a time. On our test laptop the same code turned 200 three-page PDFs into text in about 7 seconds. What can fail is a single file too big for the browser's memory, and the page says so for that file and carries on.

Why is one of my text files empty?

That PDF is a scan: its pages are pictures, so there is no text to take out. Its .txt holds 1 byte, and the result says the file has no text and to run OCR PDF first. Run it through OCR PDF, then convert it again.

Are the text files named after the PDFs?

Yes. invoice-001.pdf becomes invoice-001.txt. If 2 PDFs in different subfolders share a name, the second one's text is saved as name-2.txt, because the results land in 1 folder or zip without the subfolders.

What happens to a PDF with a password in the batch?

It is skipped and listed with the reason: it has a password. The other files are done. Run that file on its own and the page asks for its password, which stays in your tab.

Is anything uploaded?

No. The PDFs are read by an engine running in your browser tab, and the text files are made there too. After the first file, the engine is kept on your device, so a batch works with Wi-Fi off.

The tools for this job