Text copied from a PDF comes out jumbled: why, and what fixes it
By the getPDF team · Published 11 October 2026
The short answer
Pasted PDF text goes wrong in 3 ways, each with its own cause. Wrong order: a PDF stores positioned bits of text in the order its program drew them, not in reading order. Gibberish: the font has no correct map from its shapes back to letters, so only OCR can read it. Glued or spaced words: spaces are often gaps, not characters. Run the file through PDF to Text below: if that reads right, use it; if not, the cause is in the file, and the sections below say what still works.
Try it here, nothing is uploaded
Which kind of jumbled do you have?
| What you see after pasting | What causes it | What fixes it |
|---|---|---|
| Sentences from 2 columns mixed, a footer in the middle, a table read across | The order the file stores the text in | Copy 1 column at a time; or, if the file stores the right order, a converter |
Letters that are all wrong, boxes, RIILFH for “office” |
No map, or a wrong map, from font shapes back to letters | OCR, which reads the page as a picture |
| “ofce” or a strange single character where fi or fl should be | A ligature without a proper map | PDF to Text gives the letters back; or find and replace |
Rentdueonthe or R e n t |
Spaces stored as gaps the reader has to guess | Find and replace; check the guess |
| A break after every line | Lines, not paragraphs, in the file | The line break guide |
Why the order breaks
A PDF page is not a document with paragraphs. It is a list of drawing instructions: “at this position, in this font, draw these letters”. The list runs in whatever order the program that made the file wrote it. Word and most report generators write the text roughly as you read it. Some layout programs, web-to-PDF printers and form generators do not: they draw a page line by line straight across both columns, or draw the footer first, or each text frame in the order it was created.
We built both versions of a 2-column lease page and ran them through getPDF’s converters:
- Stored column by column: PDF to Text gave the left column, then the right; PDF to Word and PDF to Markdown gave 2 clean paragraphs.
- Stored line by line: every tool gave “Left column line one about Right column line one on”, then the next pair of half-lines, and so on. Word and Markdown joined those mixed lines into 1 paragraph.
- Right column drawn first: every tool put the right column first.
So be clear on what a converter can and cannot do here. getPDF’s engine, PDFium (the engine Chrome uses to show PDFs), reads the text in the order the file stores it. It does not work out columns from positions. If your paste is jumbled but PDF to Text reads correctly, the jumble came from the selection or from the viewer you copied in, and the converted text is your fix. If PDF to Text is jumbled the same way, the file stores the text in that order.
Fix the order when the file stores it mixed
Copy 1 column at a time with column selection, which selects only the text inside a rectangle you drag:
- Adobe Acrobat Reader on Windows: with the Select tool, hold Alt and drag a rectangle over the column (Adobe Help, “Reusing PDF content”, checked on 11 October 2026).
- Preview on a Mac: choose Tools, Text Selection, hold Option while you drag over the text, then Edit, Copy. Apple’s guide names copying a column of a table as the use (Apple Support, “Select and copy text in a PDF in Preview on Mac”, checked on 11 October 2026).
OCR does not reorder it either. We ran OCR PDF with Pages with text set to “read again too” on the line-by-line page; its text layer also came out line by line across both columns, at 95 % confidence. For a long document stored this way, column selection, page by page, is the honest answer; for a short one, retyping can be quicker.
Why gibberish appears
Inside a PDF, a font often does not store letters at all. It stores numbered shapes (glyphs), and the text on the page is a list of glyph numbers. To turn those numbers back into letters, the font carries a map, the ToUnicode map. Copy, search and every converter use it.
When the map is missing or wrong, what comes out is the glyph numbers read as if they were letters. We made a test file with an embedded font, removed its map, and ran it through PDF to Text. “Łódź office: the lease starts on 1 March.” came out as ăµGļRIILFHWKHOHDVHVWDUWVRQ0DUFK. Each letter is shifted by the same 29 places, because in that font the glyph for “a” happens to be number 68 while its letter code is 97: “office” became “RIILFH”. The spaces turned into invisible control characters.
No converter can recover text from a font without its map, because the information is not in the file. What does work is reading the page the way a scanner would:
- Open OCR PDF and drop the file.
- Pick the document’s language under Language.
- Set Pages with text to “read again too”. Without it, OCR leaves the page alone, because the page already has text, gibberish or not.
- Press Make it searchable, then run the result through PDF to Text.
On our test file this gave “L6dz office: the lease starts on 1 March.” at 84 % confidence, read in English: right except for the Polish place name. The gibberish stays in the file too, so the text output has the gibberish line first and the OCR line after it; delete the gibberish lines from your copy.
Inspect cannot warn you about this in advance. On the same test file it showed the font as embedded and the page as having text, which is true. The only test is to read some of the text.
Ligatures: when fi and fl go missing
Many fonts draw “fi”, “fl” and “ff” as 1 joined shape, a ligature. How it pastes depends on the map:
- If the map points the ligature at the single character “fi” (U+FB01), you paste 1 character that looks right, but it is not the letters f and i, so a plain search for “office” does not match it unless the program treats the 2 as equal.
- If there is no map entry and the shape has a name no reader knows, nothing sensible comes out. On our test file built that way, “office” came out of PDF to Text as “of”, an invisible control character, and “ce”: it reads “ofce”.
getPDF’s engine gives the letters back whenever the file holds any clue. On our test files, a font that names the shape “fi”, one that maps it to the single character, and one whose map has no entry but whose shape is named “fi” all came out of PDF to Text as “office”, and search found the word. If your paste shows the ligature character, run the file through PDF to text, or replace “fi” with “fi” and “fl” with “fl” in your editor. If letters are missing, as in “ofce”, only OCR or a careful find and replace brings them back.
Glued words and spaced letters
Some PDFs store no space characters at all: each word is placed at its own position, and the gap is the space. Whatever reads the text has to guess where words end. On our test file, 5 words placed apart without spaces came out as “Rent due on the first”, and a word placed 1 letter at a time came out as “Deposit”, with no stray spaces. The guess fails on tightly set justified text (words glued together) and on wide letter-spacing (letters spaced out). Fix it with find and replace in the destination, and read the result before you send it.
What cannot be fixed, and how to know first
Mixed order stays mixed through every converter here, OCR included, and a font without its map gives gibberish everywhere except through OCR, which guesses. Text inside pictures is not text at all; see why you cannot copy text. To know before a big job, take 1 minute with the check before converting.
Questions
Why does text from a PDF paste in the wrong order?
A PDF stores each piece of text at a position on the page, in the order the program that made it drew them. Copy and most converters follow that stored order. When a program drew a 2-column page line by line across both columns, the columns come out mixed.
Why does PDF text paste as gibberish like RIILFH?
The font in the PDF lacks a correct map from its drawn shapes back to letters (the ToUnicode map), so the shapes' internal numbers come out as letters. In our test, office became RIILFH. Only OCR, which reads the page as a picture, gets the words back.
Can a converter fix the order?
Only if the file stores the text in reading order and the jumble came from the selection. getPDF's converters, like most, keep the order the file stores; they do not reorder columns. For a file that stores columns mixed, copy 1 column at a time with column selection.
Why do fi and fl go missing or look odd when pasted?
Many fonts draw fi and fl as 1 joined shape, a ligature. If the file maps it to the single character for fi, you paste 1 character that a plain search for office does not match; if the file gives no clue at all, the letters drop out. getPDF's engine gives f and i back as 2 letters whenever it can.
The tools for this job
Keep reading
- Copied PDF text breaks at the end of every line: 3 fixesText copied from a PDF breaks at the end of every line because a PDF stores lines, not paragraphs.
- Cannot select or copy text in a PDF: the 3 causes and their fixesText in a PDF will not select or copy for 1 of 3 reasons: a scan, a copy restriction, or a font that hides the letters.
- Is the PDF a scan or real text? Check it before you convertCheck if a PDF is scanned or real text before converting: 1 minute shows text pages, scanned pages, fonts and copy rights, so you pick the right tool first.
- Make a scanned PDF searchable, without changing how it looksOCR a scanned PDF on your device: free, no account, nothing is uploaded.
- Convert many PDFs to text at once, and catch the scans that come out emptyConvert multiple PDFs to text in one batch in your browser: a whole folder in, 1 .txt per PDF out, and how to spot the scanned files that come out empty.
- PDF to Word, Excel and text: what converts, what breaks, and how to pickConvert PDF to Word, Excel, plain text, Markdown or CSV free in your browser.