Skip to content

PDF to Word, Excel and text: what converts, what breaks, and how to pick

By the getPDF team · Published 11 October 2026

The short answer

A PDF converts to an editable file on this page, free, with nothing uploaded: Word for prose you want to rewrite, Excel for tables and numbers, plain text or Markdown for notes apps and AI tools. The result is the document’s content as flowing, editable text, not a replica of the page layout, and the tool tells you that before you click. Reports, letters and contracts convert well; brochures and magazine layouts do not. A scanned PDF has no text to convert and goes through OCR first.

Pick the target format in 1 minute

The right format follows from what you do next, not from the PDF:

You want to Pick Why
Rewrite the prose, track changes, restyle Word (.docx) via PDF to Word Headings, lists, tables and pictures arrive as real Word objects you can edit
Sum, sort or chart the numbers Excel (.xlsx) via PDF to Excel Tables become cells, amounts become numbers, IDs stay text
Import into accounting or a database CSV, from the same PDF to Excel tool UTF-8 with a BOM, so accents survive; comma or semicolon separator to match your Excel
Feed a notes app or an AI tool Markdown via PDF to Markdown Headings, lists, paragraphs and tables survive as structure the tool understands
Just the words, pasted anywhere Plain text via PDF to text All the text, line by line, UTF-8, nothing else

If the document is a scan, none of these work yet; the section on scans below explains the 1 extra step.

Convert it now

Try it here, nothing is uploaded

PDF · any size · many at once

Drop the file, convert, open the result in Word. Before you click, the page says that the converter rebuilds the content, not a copy of the layout, and the result counts what was converted: headings, paragraphs, list items, tables and pictures. The .docx opens in Word, Google Docs, Pages and LibreOffice. Two options sit next to the button: running headers, footers and page numbers are left out by default, and pictures go in by default.

What a PDF actually is, and why every converter rebuilds

A PDF is not a document in the sense Word means it. It is a page description: each run of letters is placed at a position, each line is drawn where it is, each picture sits at coordinates. There are no paragraphs inside, no heading styles, no table objects, usually not even word spaces you can rely on. The file says “put these glyphs at x 72, y 410”, and that is all it says.

So no tool converts a PDF. Every tool, ours included, reconstructs: it reads positions and sizes, guesses where a paragraph ends, decides that a larger bold line is a heading, notices that aligned text with ruled lines is probably a table. Good reconstruction gets this right nearly always on documents that were born as documents. But it is detective work, and knowing that explains every quirk you will ever see in a converted file.

That is also why the choice is binary. A converter can give you editable flowing text, or it can give you an exact copy of the page, and it cannot give you both. Keep that in mind when a competitor’s page promises otherwise; the honest part below looks at what those promises hide.

What PDF to Word gives you, exactly

Our converter makes the first choice, openly: a flowing, editable .docx. In practice that means:

  • Headings become Word’s Heading 1 to Heading 3 styles. The navigation pane works, a table of contents can be generated, and restyling the whole document takes 1 click.
  • Paragraphs flow. Lines are rejoined into paragraphs, so editing a sentence reflows the text like in any normal Word file, instead of leaving a ragged hole.
  • Bullet lists become real Word lists, and numbers stay as printed. A bullet list grows a new bullet when you press Enter. Numbered items such as 3., 4.2.1 or a) keep the exact number the PDF prints, followed by a tab and a hanging indent, so “clause 4.2” in the text still points at clause 4.2; they are not an automatic Word list that renumbers itself.
  • Tables become Word tables when 3 or more rows of 3 or more cells line up (2 cells are enough when ruled lines mark the columns). Rows and columns are editable cells, not tab-separated text.
  • Pictures stay in place. JPEGs are carried over byte for byte, so photographs lose nothing in the round trip.
  • Each PDF page ends with a page break, and the page size carries over. A4 stays A4, Letter stays Letter, and content keeps its page of origin even though lines may wrap differently.

What does not carry over is the exact look: precise line breaks, multi-column flow, text wrapped tightly around images, decorative placement. That is the trade, and the tool states it before the click, not after.

Diagram of a PDF page’s positioned blocks mapping to Word structures: heading to Heading 1, ruled rows to a Word table, picture carried over.PDF page: positioned blocksQuarterly reportWord document: real structuresHeading 1 styleWord table, editable cellsPicture,byte for byte
How a PDF page's blocks map to Word structures: a large bold line becomes a Heading 1 style, aligned rows with ruled lines become an editable Word table, and a picture is carried over unchanged.

What converts well, and what does not

The document’s design decides the outcome, and you can predict it before converting:

Source document Best target What to expect Verdict
Report or thesis Word Headings, lists, tables and pictures land as Word objects; line breaks move Converts well
Letter or contract Word Clause numbers such as 4.2 or a) arrive exactly as printed, so cross references still match; check signature blocks by hand Converts well
Bank statement Excel Transactions as rows, amounts as numbers you can sum against the closing balance Converts well
Invoice Excel The line-item table becomes cells; verify the first invoice from each supplier once Converts well
Brochure or magazine spread None of these Multi-column designed pages come out as prose in the wrong order Do not convert; take the text and redesign
Scanned paper OCR first The page is a picture; conversion passes it through with a pointer to OCR Convert after OCR

The pattern: documents that were born in a word processor go back to one cleanly. Documents that were designed, where the layout is the point, do not, because the layout is exactly what flowing text gives up. For a designed page, converting the words to text and rebuilding in the destination is faster than repairing a mangled conversion. And when you are unsure whether a specific file will behave, checking it before converting takes 1 minute and tells you about its text layer and fonts; the deciding between retyping and converting also has its own honest time math for borderline cases.

PDF to Excel: numbers that stay numbers

PDF to ExcelTables into rows and columns, numbers as numbers. Free, runs on your device.

Tables are the harder half of this job, because a PDF table is only text positioned to look like a grid. The converter finds tables by position and by ruled lines, then does the part spreadsheets usually get wrong: it decides, per column, what is a number and what only looks like one.

  • Amounts written as 1,190.00 and as 1.190,00 are both read as the number 1190, so a German statement and a US invoice both sum correctly.
  • Accounting negatives in parentheses, like (12.50), become -12.50.
  • Dates are kept as text, exactly as printed, so 01.10.2026 is never turned into the wrong day by an Excel set to another country. To sort by date, convert the column once with Excel’s Data, Text to Columns, Date.
  • IDs with leading zeros, phone numbers and account numbers stay text, so Excel does not turn 0043316 into 43316 or a 16-digit card number into scientific notation.

A multi-page document gives you the choice of a sheet per page or everything on 1 sheet. For imports into accounting software there is CSV output, written as UTF-8 with a BOM, which is the encoding detail that makes Excel read names like Jana Čermák or a Łódź address correctly instead of as mojibake. Pick the Separator to match the Excel that opens it: “Comma, 1190.50” for Excel in English, Google Sheets and most software, “Semicolon, 1190,50” for Excel set to German, French, Polish and other languages that write a decimal comma; then a double-click opens the file in columns with the amounts as numbers. The .xlsx needs no such choice: it opens correctly in every language.

Two habits make table conversion trustworthy. First, verify once per source: sum the amount column of a converted statement against the printed closing balance, or an invoice’s line items against its total, 30 seconds that catches every miss. Second, expect cleanup on hard cases: a table drawn with spacing alone and no ruled lines, 2 tables sitting side by side on 1 page, or a table running across 40 pages with repeated header rows. Each of those has its own guide with fixes ordered by effort.

Plain text and Markdown, for everything else

PDF to textAll the text, in reading order, as UTF-8. Free, runs on your device.

PDF to text gives you all the text, line by line and page by page, as UTF-8, and nothing else. It is the right pick when the destination does its own formatting: a translation job, a search index, a quick paste into an email. It takes the text in the order the file stores it, which for letters, reports and most exports is the order you read in; a 2-column page whose program drew it line by line across both columns comes out with the columns mixed, so check 1 page of an unusual layout first.

PDF to Markdown adds structure to that text: heading levels, lists, paragraphs and tables as Markdown syntax, with clause numbers kept as printed. That is the format notes apps like Obsidian and Notion keep natively, and the one AI tools read best, because a model that can see headings retrieves and summarises better than one fed a wall of undifferentiated text. When you only need the words, text is enough; when the destination understands structure, Markdown is worth it.

A scanned PDF converts to a picture of itself

A scan is a photograph of a page. It contains no text at all, only pixels that look like text to you, so there is nothing for any converter to extract. Put a scan through PDF to Word here and the page comes out as its picture, with a pointer telling you why and where to go.

You can spot the case in 5 seconds without any tool: try to select a sentence in your PDF viewer. If nothing selects, or the selection is 1 big box around the whole page, and the letters go blurry when you zoom in, it is a scan.

The route is OCR first: it reads the pixels like you do and writes a real text layer into the file, after which every conversion on this page works normally. The complete guide to that step, including languages and accuracy, is Scanned PDF to searchable. Mixed files exist too: a digital contract with 1 scanned signature page converts fine except for that page.

The honest part: pixel-perfect, free and private cannot all be true

Search for “pdf to word” and the results promise perfect conversion, exact layout, just like the original. Here is what those promises are made of, because a PDF can only become an editable Word file in 2 ways:

  1. Flowing text, which is what we build: real paragraphs, styles, lists and tables that edit like a document written in Word, at the cost of exact page layout.
  2. A layout replica, built by placing hundreds of absolutely positioned text boxes, or whole page images, exactly where the PDF had them. It looks identical in the first screenshot and edits terribly: text does not reflow, boxes overlap when you type, and restyling is hopeless.

Converters that deliver genuinely good-looking and reasonably editable results do the heavy reconstruction on a server, which means your contract, statement or HR file is uploaded to someone else’s machine, usually under a free tier with daily limits designed to sell a subscription. That is the whole trade-off triangle: pixel-perfect, editable and on-your-device; pick 2. We picked editable and on-your-device, and we label the cost before you click.

There is also a case where the right answer is not converting at all. If you need to change a date, a name or a paragraph and the document must keep looking exactly as it does, edit the PDF directly and leave the layout untouched; the Word round trip is a detour that costs you the formatting for no gain. Edit a PDF covers that path.

Nothing is uploaded, and you can check

The converters on this page run on your device, in your browser. The engine (4.6 MB) downloads from this site the first time you drop a file, and your browser keeps it; your file stays in your browser’s memory, and the result is written straight to your disk. No account, no file-size tier, no daily limit, no watermark.

You do not have to take that on trust. Convert one file, turn Wi-Fi off, then drop another and convert: it works, because there is nothing left to talk to. The site’s security policy enforces the same thing technically, by forbidding the page from sending data anywhere else. For bank statements, contracts and anything with personal data in it, that is the property that matters most, and it is the reason this converter exists. The wider story, including what happened when online converter sites leaked user documents, is in PDF privacy and protection.

Questions

Does the Word file look exactly like the PDF?

No, and no honest converter's does. The output is the document's content as flowing, editable text with headings, lists, tables and pictures, not a replica of the page. Converters that show a pixel-perfect copy get it by stacking positioned text boxes, which fight every edit, or by processing your file on a server.

Why did my PDF convert to a picture instead of text?

The file is a scan. A scanned page is a photograph of text with no text in it, so the converter can only pass the picture through and point you to OCR. Run the file through OCR first to give it a real text layer, then convert.

Does PDF to Excel keep leading zeros and long account numbers?

Yes. The converter reads amounts like 1,190.00 or 1.190,00 as numbers, but keeps dates, IDs with leading zeros, phone numbers and account numbers as text, so Excel does not strip the zeros or switch to scientific notation.

Is my file uploaded to convert it?

No. The converter runs in your browser, on your device. The only network traffic is the page and the engine, fetched from this site when you drop your first file and kept after that. You can check it: convert once, turn Wi-Fi off, and the next conversion still works.

Which format should I pick for an AI tool or a notes app?

Markdown. It keeps headings, lists, paragraphs and tables as structure the tool can read, which beats raw text for retrieval and summarising. When only the words matter, plain text is enough and smaller.

The tools for this job