PDF to Word, Excel and text: what converts, what breaks, and how to pick
By the getPDF team · Published 11 October 2026
The short answer
A PDF converts to an editable file on this page, free, with nothing uploaded: Word for prose you want to rewrite, Excel for tables and numbers, plain text or Markdown for notes apps and AI tools. The result is the document’s content as flowing, editable text, not a replica of the page layout, and the tool tells you that before you click. Reports, letters and contracts convert well; brochures and magazine layouts do not. A scanned PDF has no text to convert and goes through OCR first.
Pick the target format in 1 minute
The right format follows from what you do next, not from the PDF:
| You want to | Pick | Why |
|---|---|---|
| Rewrite the prose, track changes, restyle | Word (.docx) via PDF to Word | Headings, lists, tables and pictures arrive as real Word objects you can edit |
| Sum, sort or chart the numbers | Excel (.xlsx) via PDF to Excel | Tables become cells, amounts become numbers, IDs stay text |
| Import into accounting or a database | CSV, from the same PDF to Excel tool | UTF-8 with a BOM, so accents survive; comma or semicolon separator to match your Excel |
| Feed a notes app or an AI tool | Markdown via PDF to Markdown | Headings, lists, paragraphs and tables survive as structure the tool understands |
| Just the words, pasted anywhere | Plain text via PDF to text | All the text, line by line, UTF-8, nothing else |
If the document is a scan, none of these work yet; the section on scans below explains the 1 extra step.
Convert it now
Try it here, nothing is uploaded
Drop the file, convert, open the result in Word. Before you click, the page says that the converter rebuilds the content, not a copy of the layout, and the result counts what was converted: headings, paragraphs, list items, tables and pictures. The .docx opens in Word, Google Docs, Pages and LibreOffice. Two options sit next to the button: running headers, footers and page numbers are left out by default, and pictures go in by default.
What a PDF actually is, and why every converter rebuilds
A PDF is not a document in the sense Word means it. It is a page description: each run of letters is placed at a position, each line is drawn where it is, each picture sits at coordinates. There are no paragraphs inside, no heading styles, no table objects, usually not even word spaces you can rely on. The file says “put these glyphs at x 72, y 410”, and that is all it says.
So no tool converts a PDF. Every tool, ours included, reconstructs: it reads positions and sizes, guesses where a paragraph ends, decides that a larger bold line is a heading, notices that aligned text with ruled lines is probably a table. Good reconstruction gets this right nearly always on documents that were born as documents. But it is detective work, and knowing that explains every quirk you will ever see in a converted file.
That is also why the choice is binary. A converter can give you editable flowing text, or it can give you an exact copy of the page, and it cannot give you both. Keep that in mind when a competitor’s page promises otherwise; the honest part below looks at what those promises hide.
What PDF to Word gives you, exactly
Our converter makes the first choice, openly: a flowing, editable .docx. In practice that means:
- Headings become Word’s Heading 1 to Heading 3 styles. The navigation pane works, a table of contents can be generated, and restyling the whole document takes 1 click.
- Paragraphs flow. Lines are rejoined into paragraphs, so editing a sentence reflows the text like in any normal Word file, instead of leaving a ragged hole.
- Bullet lists become real Word lists, and numbers stay as printed. A bullet list grows a new bullet when you press Enter. Numbered items such as 3., 4.2.1 or a) keep the exact number the PDF prints, followed by a tab and a hanging indent, so “clause 4.2” in the text still points at clause 4.2; they are not an automatic Word list that renumbers itself.
- Tables become Word tables when 3 or more rows of 3 or more cells line up (2 cells are enough when ruled lines mark the columns). Rows and columns are editable cells, not tab-separated text.
- Pictures stay in place. JPEGs are carried over byte for byte, so photographs lose nothing in the round trip.
- Each PDF page ends with a page break, and the page size carries over. A4 stays A4, Letter stays Letter, and content keeps its page of origin even though lines may wrap differently.
What does not carry over is the exact look: precise line breaks, multi-column flow, text wrapped tightly around images, decorative placement. That is the trade, and the tool states it before the click, not after.
What converts well, and what does not
The document’s design decides the outcome, and you can predict it before converting:
| Source document | Best target | What to expect | Verdict |
|---|---|---|---|
| Report or thesis | Word | Headings, lists, tables and pictures land as Word objects; line breaks move | Converts well |
| Letter or contract | Word | Clause numbers such as 4.2 or a) arrive exactly as printed, so cross references still match; check signature blocks by hand | Converts well |
| Bank statement | Excel | Transactions as rows, amounts as numbers you can sum against the closing balance | Converts well |
| Invoice | Excel | The line-item table becomes cells; verify the first invoice from each supplier once | Converts well |
| Brochure or magazine spread | None of these | Multi-column designed pages come out as prose in the wrong order | Do not convert; take the text and redesign |
| Scanned paper | OCR first | The page is a picture; conversion passes it through with a pointer to OCR | Convert after OCR |
The pattern: documents that were born in a word processor go back to one cleanly. Documents that were designed, where the layout is the point, do not, because the layout is exactly what flowing text gives up. For a designed page, converting the words to text and rebuilding in the destination is faster than repairing a mangled conversion. And when you are unsure whether a specific file will behave, checking it before converting takes 1 minute and tells you about its text layer and fonts; the deciding between retyping and converting also has its own honest time math for borderline cases.
PDF to Excel: numbers that stay numbers
PDF to ExcelTables into rows and columns, numbers as numbers. Free, runs on your device.Tables are the harder half of this job, because a PDF table is only text positioned to look like a grid. The converter finds tables by position and by ruled lines, then does the part spreadsheets usually get wrong: it decides, per column, what is a number and what only looks like one.
- Amounts written as 1,190.00 and as 1.190,00 are both read as the number 1190, so a German statement and a US invoice both sum correctly.
- Accounting negatives in parentheses, like (12.50), become -12.50.
- Dates are kept as text, exactly as printed, so 01.10.2026 is never turned into the wrong day by an Excel set to another country. To sort by date, convert the column once with Excel’s Data, Text to Columns, Date.
- IDs with leading zeros, phone numbers and account numbers stay text, so Excel does not turn 0043316 into 43316 or a 16-digit card number into scientific notation.
A multi-page document gives you the choice of a sheet per page or everything on 1 sheet. For imports into accounting software there is CSV output, written as UTF-8 with a BOM, which is the encoding detail that makes Excel read names like Jana Čermák or a Łódź address correctly instead of as mojibake. Pick the Separator to match the Excel that opens it: “Comma, 1190.50” for Excel in English, Google Sheets and most software, “Semicolon, 1190,50” for Excel set to German, French, Polish and other languages that write a decimal comma; then a double-click opens the file in columns with the amounts as numbers. The .xlsx needs no such choice: it opens correctly in every language.
Two habits make table conversion trustworthy. First, verify once per source: sum the amount column of a converted statement against the printed closing balance, or an invoice’s line items against its total, 30 seconds that catches every miss. Second, expect cleanup on hard cases: a table drawn with spacing alone and no ruled lines, 2 tables sitting side by side on 1 page, or a table running across 40 pages with repeated header rows. Each of those has its own guide with fixes ordered by effort.
Plain text and Markdown, for everything else
PDF to textAll the text, in reading order, as UTF-8. Free, runs on your device.PDF to text gives you all the text, line by line and page by page, as UTF-8, and nothing else. It is the right pick when the destination does its own formatting: a translation job, a search index, a quick paste into an email. It takes the text in the order the file stores it, which for letters, reports and most exports is the order you read in; a 2-column page whose program drew it line by line across both columns comes out with the columns mixed, so check 1 page of an unusual layout first.
PDF to Markdown adds structure to that text: heading levels, lists, paragraphs and tables as Markdown syntax, with clause numbers kept as printed. That is the format notes apps like Obsidian and Notion keep natively, and the one AI tools read best, because a model that can see headings retrieves and summarises better than one fed a wall of undifferentiated text. When you only need the words, text is enough; when the destination understands structure, Markdown is worth it.
A scanned PDF converts to a picture of itself
A scan is a photograph of a page. It contains no text at all, only pixels that look like text to you, so there is nothing for any converter to extract. Put a scan through PDF to Word here and the page comes out as its picture, with a pointer telling you why and where to go.
You can spot the case in 5 seconds without any tool: try to select a sentence in your PDF viewer. If nothing selects, or the selection is 1 big box around the whole page, and the letters go blurry when you zoom in, it is a scan.
The route is OCR first: it reads the pixels like you do and writes a real text layer into the file, after which every conversion on this page works normally. The complete guide to that step, including languages and accuracy, is Scanned PDF to searchable. Mixed files exist too: a digital contract with 1 scanned signature page converts fine except for that page.
The honest part: pixel-perfect, free and private cannot all be true
Search for “pdf to word” and the results promise perfect conversion, exact layout, just like the original. Here is what those promises are made of, because a PDF can only become an editable Word file in 2 ways:
- Flowing text, which is what we build: real paragraphs, styles, lists and tables that edit like a document written in Word, at the cost of exact page layout.
- A layout replica, built by placing hundreds of absolutely positioned text boxes, or whole page images, exactly where the PDF had them. It looks identical in the first screenshot and edits terribly: text does not reflow, boxes overlap when you type, and restyling is hopeless.
Converters that deliver genuinely good-looking and reasonably editable results do the heavy reconstruction on a server, which means your contract, statement or HR file is uploaded to someone else’s machine, usually under a free tier with daily limits designed to sell a subscription. That is the whole trade-off triangle: pixel-perfect, editable and on-your-device; pick 2. We picked editable and on-your-device, and we label the cost before you click.
There is also a case where the right answer is not converting at all. If you need to change a date, a name or a paragraph and the document must keep looking exactly as it does, edit the PDF directly and leave the layout untouched; the Word round trip is a detour that costs you the formatting for no gain. Edit a PDF covers that path.
Nothing is uploaded, and you can check
The converters on this page run on your device, in your browser. The engine (4.6 MB) downloads from this site the first time you drop a file, and your browser keeps it; your file stays in your browser’s memory, and the result is written straight to your disk. No account, no file-size tier, no daily limit, no watermark.
You do not have to take that on trust. Convert one file, turn Wi-Fi off, then drop another and convert: it works, because there is nothing left to talk to. The site’s security policy enforces the same thing technically, by forbidding the page from sending data anywhere else. For bank statements, contracts and anything with personal data in it, that is the property that matters most, and it is the reason this converter exists. The wider story, including what happened when online converter sites leaked user documents, is in PDF privacy and protection.
Questions
Does the Word file look exactly like the PDF?
No, and no honest converter's does. The output is the document's content as flowing, editable text with headings, lists, tables and pictures, not a replica of the page. Converters that show a pixel-perfect copy get it by stacking positioned text boxes, which fight every edit, or by processing your file on a server.
Why did my PDF convert to a picture instead of text?
The file is a scan. A scanned page is a photograph of text with no text in it, so the converter can only pass the picture through and point you to OCR. Run the file through OCR first to give it a real text layer, then convert.
Does PDF to Excel keep leading zeros and long account numbers?
Yes. The converter reads amounts like 1,190.00 or 1.190,00 as numbers, but keeps dates, IDs with leading zeros, phone numbers and account numbers as text, so Excel does not strip the zeros or switch to scientific notation.
Is my file uploaded to convert it?
No. The converter runs in your browser, on your device. The only network traffic is the page and the engine, fetched from this site when you drop your first file and kept after that. You can check it: convert once, turn Wi-Fi off, and the next conversion still works.
Which format should I pick for an AI tool or a notes app?
Markdown. It keeps headings, lists, paragraphs and tables as structure the tool can read, which beats raw text for retrieval and summarising. When only the words matter, plain text is enough and smaller.
The tools for this job
Every guide in PDF to Word and Excel
- Convert a PDF to Word you can edit, not a page of stuck text boxesConvert a PDF to a Word file you can really edit: flowing paragraphs, Word headings, lists and tables, not stacked text boxes.
- Bank statement to Excel, without uploading it anywhereTurn a bank statement PDF into an Excel sheet without uploading it anywhere: steps, amount checks, and what to do when the statement is a scanned image.
- PDF to Markdown for notes and AI tools: what carries over and what flattensConvert a PDF to Markdown for Obsidian, Notion or an AI assistant: headings, lists and tables carry over; columns and emphasis flatten.
- Retype or convert a PDF? The honest time mathSometimes retyping a PDF beats converting it: a short letter, a designed page, a tiny table.
- Is the PDF a scan or real text? Check it before you convertCheck if a PDF is scanned or real text before converting: 1 minute shows text pages, scanned pages, fonts and copy rights, so you pick the right tool first.
- Make a scanned PDF searchable, without changing how it looksOCR a scanned PDF on your device: free, no account, nothing is uploaded.
- Convert many PDFs to text at once, and catch the scans that come out emptyConvert multiple PDFs to text in one batch in your browser: a whole folder in, 1 .txt per PDF out, and how to spot the scanned files that come out empty.
- Cannot select or copy text in a PDF: the 3 causes and their fixesText in a PDF will not select or copy for 1 of 3 reasons: a scan, a copy restriction, or a font that hides the letters.
- Contract PDF to Word for tracked changes, with the clause numbers checkedTurn a contract PDF into Word for tracked changes: convert on your device, clause numbers kept as printed, check tables and amounts, then redline and send.
- Text copied from a PDF comes out jumbled: why, and what fixes itCopy and paste from a PDF comes out jumbled: wrong order, gibberish or glued words.
- Copy a table from a PDF without the columns collapsingCopy a table from a PDF into Excel or Word: why paste lands everything in 1 column, when splitting it works, and the 2-minute convert route when it does not.
- Invoice PDFs into Excel for bookkeeping, checked against the totalGet invoice line items from a PDF into Excel: amounts as numbers, a 1-minute check against the printed total, and when the e-invoice XML is the better source.
- A table across many PDF pages, into 1 clean Excel sheetTurn a table that spans many PDF pages into 1 Excel sheet: repeated headings and footers out, broken rows mended, the row count checked.
- Copied PDF text breaks at the end of every line: 3 fixesText copied from a PDF breaks at the end of every line because a PDF stores lines, not paragraphs.
- Get a table from a PDF into Google Sheets, with the numbers intactGet a table from a PDF into Google Sheets: convert it to .xlsx on your device, then File, Import.
- PDF to CSV your accounting software actually importsTurn a PDF statement or invoice into a CSV your accounting software imports: convert on your device, tidy the rows, save as CSV UTF-8, check in a text editor.
- Leading zeros disappear in Excel: account, customer and phone numbers, fixedExcel turns 004711 into 4711 and card numbers into 4.11E+15 when it reads codes as numbers.
- Amounts and dates come out wrong in Excel: decimal commas and localesAmounts like 1.234,56 turn into text and 03/04 into the wrong date when Excel reads them in another locale.
- PDF to Excel puts text in the wrong columns: why, and the fixesPDF to Excel put everything in 1 column, or 2 columns in 1 cell? How converters find columns, the 4 layouts that defeat them, and a fix for each.
- PDF to Google Docs: the 2 ways that work, and which keeps your tablesOpen a PDF in Google Docs 2 ways: Drive's own import, or convert to Word first.
- Get the text out of a PDF for translation, clean enough to quote and translateExtract text from a PDF for translation: a Word file with whole paragraphs, no page numbers or headers in the flow, and an honest word count for the quote.
- The converted Word file looks different: what changed, and the 5-minute fixesYour PDF to Word file looks different because page layout does not survive conversion.
- Convert a PDF to Word offline: no upload, and you can prove itConvert a PDF to Word with nothing uploaded: the converter runs in your browser and keeps working with Wi-Fi off.
- Convert a PDF to Word on your phone, with no app to installConvert a PDF to Word on iPhone or Android in the browser: no app, nothing uploaded.