Skip to content

Copied PDF text breaks at the end of every line: 3 fixes

By the getPDF team · Published 11 October 2026

The short answer

The line breaks are in the PDF: it stores printed lines placed one under the other, not paragraphs, and copy gives you those lines. For a paragraph or two, paste into Word and replace the paragraph marks (^p) with spaces. For a whole document, convert it with PDF to Word below, which rebuilds paragraphs from the line spacing. For plain text with whole paragraphs, use PDF to Markdown. Then proofread the words that were hyphenated at the margin.

Try it here, nothing is uploaded

PDF · any size · many at once

Why it happens

A PDF page holds text the way a printing plate does: this line at this height, the next line 14 points lower. There is no paragraph object and no “this line continues on the next” flag. Word wraps lines as you type; a PDF has the lines already wrapped and frozen. Copy takes the lines as they are and ends each with a break, so a 5-line paragraph pastes as 5 short paragraphs.

Fix 1, for a few paragraphs: find and replace in Word

  1. Paste the text into Word on Windows or Mac.
  2. Select the paragraphs you want to repair.
  3. Open Find and Replace (Ctrl+H on Windows).
  4. In Find what type ^p (the paragraph mark; it must be a lower-case p), in Replace with type a single space.
  5. Choose Replace All. If Word offers to search the rest of the document, answer No.

If Word finds nothing, the breaks are manual line breaks instead: search for ^l (a lower-case L). Make sure Use wildcards under More is switched off, or ^p is not recognised. These codes work in the Word desktop app; you can also pick them from the Special button instead of typing them (Word MVP site, “Finding and replacing non-printing characters”, checked on 11 October 2026).

The double-break trick for longer text. If your paste has an empty line between paragraphs, keep the real paragraph ends like this: replace ^p^p with a marker that appears nowhere else, such as ###; replace ^p with a space; then replace ### with ^p. This only works when the paste has those empty lines. Text copied from many PDFs has none, because the gap between paragraphs is just a slightly larger distance on the page; then the trick cannot see where paragraphs end, and Fix 2 or 3 does.

Fix 2, for a document: PDF to Word

PDF to Word does not copy lines. It reads each line’s position and size and decides where paragraphs start: a gap of more than about half a line, a change in font size, a switch between bold and regular, a bullet or a number starts a new one. Headings come from text clearly larger than the body; running headers, footers and page numbers are left out by default (Running lines: “left out”).

We ran it on a test paragraph of 4 printed lines, followed by a separate 1-line paragraph:

Above, plain text with a hard break after each printed line. Below, the same words as 1 flowing paragraph, then the second paragraph on its own. The word except, split at the margin as ex and cept, is joined in both.PDF to Text: 1 line per printed lineThe tenant pays the rent on the first working dayof each month. Any repairs over 200 euros need thelandlord’s written approval before work starts, except in …A second paragraph starts here and ends here.(no empty line between the 2 paragraphs)PDF to Word and PDF to Markdown: paragraphsThe tenant pays the rent on the first working day of each month.Any repairs over 200 euros need the landlord’s written approvalbefore work starts, except in an emergency such as a burst pipe.A second paragraph starts here and ends here.
4 lines of a test lease as PDF to Text keeps them (above) and as PDF to Word and PDF to Markdown write them (below). The break at the hyphen is joined in all of them.

The result is a Word file whose paragraphs wrap freely, with nothing left to repair. Copy from it into wherever the text is going.

Fix 3, for plain text: PDF to Markdown

When the destination is plain text (an email, a form field, a machine translation box, a notes app), use PDF to Markdown. It applies the same paragraph rules as PDF to Word and writes each paragraph as 1 line with an empty line between paragraphs. On the test file above it gave exactly the 2 paragraphs shown. Headings start with # and list items with -; delete those marks if the destination does not want them.

PDF to text does not do this, on purpose. It keeps 1 line per printed line, page by page, which is what you want for search indexes, line-by-line comparisons and quoting a page exactly. On our test file it kept the 4 lines and put no empty line between the 2 paragraphs.

The hyphen trap

A word split at the margin (“docu-” at the end of one line, “ment” at the start of the next) has to be joined, and the hyphen dropped. A word that really contains a hyphen (“well-known”) has to be joined and the hyphen kept. The file usually does not say which is which, so every automatic fix guesses.

On our test files, getPDF’s engine joined “ex-” and “cept” into “except”, correctly, in PDF to Text, PDF to Markdown and PDF to Word alike. It also turned “well-” at a line end followed by “known” into “wellknown”, which is wrong. Replacing ^p with a space in Word has the opposite problem: the split word becomes “ex- cept”, hyphen and space included.

So after any fix, look at the words that sat at the right margin in the PDF. A quick way: search the PDF for the hyphenated words you can see at line ends and check each in your text. In Word, searching your text for - (a hyphen followed by a space) finds the splits find and replace left behind.

When the lines are not even in the right order

If the pasted lines are complete but jumbled, columns mixed or a footer in the middle of a sentence, that is a different problem: the order the file stores its text in. The guide to jumbled PDF text explains it and what can fix it. If nothing selects at all, see why you cannot copy text from a PDF.

What these fixes cannot do

Paragraph detection works from spacing, so it can be fooled. A paragraph that continues at the top of the next page starts a new paragraph in the output; a block of short lines with tight spacing, like an address, can be joined into 1 line; text in 2 columns stored line by line comes out mixed, whatever the fix. PDF to Word and PDF to Markdown both ask you in their result to check the structure, and that is the honest instruction: read the text once before you send it on.

Questions

Why does text from a PDF have a line break after every line?

Because the PDF has no paragraphs. It stores each printed line as text placed at a position, one under the other. Copy gives you exactly those lines, each ending in a break.

What is the quickest fix for 1 or 2 paragraphs?

Paste into Word, select the paragraph, open Find and Replace (Ctrl+H), find ^p, replace with a space, and choose Replace All for the selection only. Then check the words that were split by a hyphen at the margin.

Does PDF to Text remove the line breaks?

No, by design: it keeps 1 line per printed line, like the page. It does join a word split by a hyphen at the line end. For whole paragraphs as plain text, use PDF to Markdown; for a document, PDF to Word.

Why did well-known turn into wellknown?

Every automatic fix has to decide whether a hyphen at the end of a line is a break in the word or part of it. On our test file, the engine joined ex-cept into except, correctly, and made well-known into wellknown. Proofread words that sat at the margin.

The tools for this job