Skip to content

Get the text out of a PDF for translation, clean enough to quote and translate

By the getPDF team · Published 11 October 2026

The short answer

First ask for the original Word or design file: translators work best from it. If the PDF is all you have, convert it with PDF to Word below: you get whole paragraphs instead of 1 line per printed line, headings kept, and the running headers and page numbers left out of the text. Check the numbers and names against the PDF, then read the word count in Word’s status bar for the quote. A scanned PDF needs OCR first. Nothing is uploaded.

Try it here, nothing is uploaded

PDF · any size · many at once

What a translator needs from you

Editable text with its paragraphs intact, in a file their software opens. Many translators work in a CAT tool (computer-assisted translation software) that splits the text into sentences, remembers earlier translations and puts the translation back into the same file format.

memoQ, one of the common CAT tools, states the trade plainly in its documentation (checked on 11 October 2026): it imports a PDF by converting it to a Word file first, or as plain text, which loses all formatting and is not recommended; it cannot export a PDF; it cannot import a password-protected PDF; and it does not read text from scanned pages. Its advice is to translate the original document where possible. So a Word file you made and checked yourself is a better thing to send than a PDF the translator’s software has to guess at.

The workflow, step by step

  1. Check the PDF in 1 minute. Drop it on Inspect. “Pages with text” should match the page count; “Scanned pages (image, no text)” should be 0; “Copy allowed” should say yes. If any of these is off, the check before converting guide says what to do.
  2. Convert with PDF to Word. In the tool above, leave Running lines at “left out” and decide on Pictures: “put in” keeps photos and logos where they sit, “left out” gives a lighter file of text only. Press Convert to Word.
  3. Read what was done. The result names it, for example “Wrote 4 headings, 12 paragraphs, 0 list items, 0 tables and 0 pictures as editable Word content”.
  4. Check the numbers, names and dates against the PDF, side by side. Amounts, article numbers, addresses and spellings like “Čermák” are where a wrong character costs the most once it is translated.
  5. Count the words in the converted file (next section).
  6. Send the .docx, with the PDF as a reference for how it looked.

The cleanup the converter does for you

Raw text out of a PDF is the printed page line by line: every line ends in a break, a word split at the margin stays split, and the header and page number of every page sit in the middle of the text. A translator’s software treats each of those breaks as the end of a segment, so 1 sentence becomes 3 half-sentences to translate.

Above: 5 lines of raw text, starting with the running header and Page 1 of 4, then lease lines that break mid-sentence. Below: the same text as 1 heading and 1 continuous paragraph.plain text, as printedResidential lease agreement, Hauptplatz 4, LinzPage 1 of 4Section 1: PartiesThis lease is made between Jane Cooper, the landlord, andJana Čermák, the tenant, for the flat atPDF to WordSection 1: Parties (a heading)This lease is made between Jane Cooper, the landlord, and JanaČermák, the tenant, for the flat at Hauptplatz 4 in Linz. Thetenancy starts on 1 March 2027 … (1 paragraph, wrapping freely)header and page number left out
The same lease page, as plain text (above) and as PDF to Word writes it (below). The header and page number are gone, the lines are 1 paragraph again, and the word split at the margin is whole.

On our invented 4-page lease, 60 printed lines of text became 12 paragraphs and 4 headings in Word. The header and “Page 1 of 4” lines were left out on every page, and “land-” at the end of one line and “lord” at the start of the next came back as “landlord”. Clause numbers such as 4.2.1 or (a) come through as printed, each followed by a tab, so the translation keeps the contract’s own numbering.

What is left for you: tables (check that cells landed in the right columns), footnotes (they come out as ordinary text at the end of their page), and any text inside pictures, which is not text at all and is not extracted. Read the line break guide for the one join no converter gets right every time: words that carry a real hyphen at the end of a line.

The word count for the quote

If your translator prices by the word, count before you ask for prices. Count the converted Word file, not text copied straight from the PDF. Word shows the count in the status bar at the bottom of the window, and Review > Word Count opens the full figures (Microsoft Support, “Show word count”, checked on 11 October 2026). If the status bar shows no count, right-click the status bar and tick Word Count.

Why not count the raw text? On the 4-page lease, raw text out of PDF to text counted 1,132 words; the Word file 1,092. The 40 extra words are the running header and page number, 10 words a page. On a 40-page contract that would be about 400 words you pay to have translated and then delete.

If the translator quotes per character instead, the PDF to text result line gives the characters directly, for example “Took 21,891 characters of text from 500 pages, in the order a PDF reader finds it, as UTF-8”. That count includes headers and page numbers, so treat it as an upper bound.

When plain text is what they asked for

Some translators, and every machine translation box, want plain text. PDF to text gives you all the text as UTF-8, but it keeps 1 line per printed line, and the headers and page numbers stay in. For text with whole paragraphs, use PDF to Markdown instead: it joins the lines into paragraphs, leaves out the running lines, and marks headings with # and lists with -, which you can delete or leave in. Paste the paragraphs, not the lines, and machine translation stops breaking sentences in half.

Layout is the translator’s last step, not yours

A copy of the page layout with the source text in it is not what you want to send. The translated text rarely fits the same boxes, so the layout has to be redone around the translation anyway, by the translator or a designer, in the original program or in Word. What you can do is keep the PDF as the reference for how it should look in the end.

What this cannot do

A scanned contract or certificate has no text to extract. Run it through OCR PDF first, then convert. OCR guesses and reports its confidence per page; names, stamps and numbers on a scan deserve a careful check, and handwriting is weaker still. The scan guide has the details.

The converter takes text in the order the file stores it. Most contracts and letters store it in reading order; a 2-column layout drawn line by line across the page comes out with the columns mixed, and no setting here reorders it. Check 1 page of a complicated layout before sending the whole file.

Finally, the privacy part. Contracts, medical letters, birth and marriage certificates are exactly the files that are sent for translation, and exactly the files not to put through a converter that uploads them. Here the conversion runs in your browser tab and nothing is uploaded; the privacy guide explains how to check that yourself. Whether a certified translation needs the original paper or a scan is a question for the translator or the office receiving it; this guide does not give legal advice.

Questions

Which format should I send to a translator?

The original Word or design file if you have it. If you only have the PDF, a Word file made from it, checked by you. Translation software opens .docx well; memoQ's own documentation calls plain text from a PDF a route that loses all formatting and is not recommended.

How do I count the words for a quote?

Convert to Word, open the file, and read the count in Word's status bar at the bottom of the window. Count the converted file, not raw copied text: on our 4-page test lease, raw text counted 40 words more, all of them repeated headers and page numbers.

My PDF is a scan. Can I still get the text?

Not directly: a scan is pictures of pages, with no text to take out. Run it through OCR PDF first, then convert. Check names and numbers carefully afterwards, because OCR reads by guessing and says how sure it is.

Is it safe to convert a contract here?

The conversion runs in your browser tab and the file is not uploaded, so the contract does not pass through anyone's server. You can check: convert 1 file, turn Wi-Fi off, and the next conversion still works.

The tools for this job