Skip to content

PDF to Markdown for notes and AI tools: what carries over and what flattens

By the getPDF team · Published 11 October 2026

The short answer

Drop the PDF on the converter below and press Convert to Markdown: you get a .md file with headings (up to 3 levels, worked out from font sizes), paragraphs joined from their lines, bulleted and numbered lists with clause numbers as printed, and tables as Markdown tables, with running headers, footers and page numbers left out. Multi-column layouts, pictures and bold or italic text do not carry over cleanly, so check the first file before converting 50. It runs on your device; the PDF is not uploaded.

Try it here, nothing is uploaded

PDF · any size · many at once

Why Markdown

Markdown is plain text with a few marks for structure: # for a heading, - for a list item. That makes it the format notes apps such as Obsidian are built on, it stays readable in any text editor forever, and version control can show what changed line by line. For AI assistants it has a further use: the headings tell the model where a section starts and what it is about, which raw text loses.

What you give up is everything about how the PDF looked: fonts, positions, page layout, colours. Markdown keeps the words and their structure, nothing else.

What carries over, and what flattens

We built a 4-page test report with the usual parts (a title, numbered headings, a bulleted and a numbered list, a table, a footnote, a 2-column passage, a web address, an italic sentence, and a running header and page number on every page) and converted it with the default settings. The result line counted the headings, paragraphs and list items it found and said “Left out 8 running lines (headers, footers, page numbers)”. We checked the table and the clause numbers again on 11 October 2026, after the converter learned to write tables.

In the PDF In the Markdown
title in large type # Quarterly Report
section headings in 2 sizes ## 1. Summary, ### 1.1 Results by region
bulleted list - Revenue up 12 %, 1 line per item
numbered list 1. Hire in Graz, numbers kept
a paragraph over several lines joined into 1 paragraph; a word split by a hyphen at a line end joined again
running header and page number left out, on all 4 pages
web address kept as plain text, not as a Markdown link
italic sentence, bold label plain text in our test; the emphasis is lost
clause numbers such as 4.2.1 or (a) kept exactly as printed, 1 clause per line
a 4-row table a Markdown table: a header row (Region, Q2, Q3), then 1 row per line
2 columns side by side lines from the 2 columns mixed into 1 paragraph
footnote a paragraph where it sits on the page, at the bottom
pictures left out

Here is the start of the file it wrote, unedited:

# Quarterly Report

## 1. Summary

Sales rose in every region. See https://example.com/q3 for the data. This sentence is in italics for emphasis.

Key points

- Revenue up 12 %
- Two new stores
- Costs flat
1. Hire in Graz
2. Open in Linz
3. Review prices

### 1.1 Results by region

| Region | Q2 | Q3 |
| --- | --- | --- |
| Vienna | 1,200 | 1,350 |
| Graz | 800 | 910 |
| Linz | 640 | 700 |

Headings, lists and tables are the part that works well, and they are the part that matters most for notes and for AI tools. Clause numbers come through as printed: a contract’s 4.2.1 stays 4.2.1 and (a) stays (a), each clause on its own line, instead of being renumbered as a Markdown list. The converter puts no blank line between them, so a Markdown viewer that follows the strict rules shows a run of clauses as 1 paragraph; add a blank line between clauses if yours does. The structure is worked out from how the text looks, so a large pull quote can turn into a heading; the tool says so in its result (“check the structure”).

The column result depends on the file. Our test file writes its 2 columns line by line, left, right, left, right, which is what the converter followed. Many PDFs write each column in one go and come out in the right order, but you only know by looking.

Steps: from PDF to notes

  1. Drop the PDF on the converter above (several at once is fine; each gets its own .md).
  2. Running lines: leave on “left out” to drop headers, footers and page numbers that repeat on most pages. Set “kept” if the header carries something you need.
  3. Page markers: switch on if you want to find your way back to the PDF. Each page then starts with <!-- page 2 -->, a comment that Markdown viewers hide.
  4. Press Convert to Markdown and download. report.pdf becomes report.md.
  5. Read the result line and the warnings. It names pages with no text (“has no text on pages 3, 4: probably scanned”) and reminds you that headings, paragraphs and tables are worked out from font sizes, spacing and positions, so check the structure.

Then into the place it goes:

  • Obsidian. A vault is a folder of Markdown files. Obsidian’s help says you can drag the file into the File explorer pane, or move it into the vault folder with Explorer or Finder.
  • Notion. Notion’s help gives Settings, Import, Text & Markdown, then the .md file. It notes a 5 MB limit per file on the Free plan and that import works on desktop and web, not in the mobile app. Its help lists headings and lists among what imports well.
  • An AI assistant. Open the .md in a text editor, copy the sections you need, and paste them. The headings help the model keep sections apart.

The Obsidian and Notion steps follow their own help pages, checked on 11 October 2026.

For AI use: Markdown or plain text?

Pick Markdown when structure matters: long documents you will ask questions about, where “what does section 4 say” needs a section 4 to exist, or a set of documents you will split by heading. Pick PDF to text when the text is all you need, for example to count words, translate, or search, or when the PDF is a single block of prose with no headings to find.

PDF to textEvery page's text, line by line, as a .txt file. Free, runs on your device.

Converting here also means you decide what leaves your device. The PDF is read in your browser; nothing goes anywhere until you paste text into an assistant or import a file into a cloud notes app, and then only the text you chose. A contract or a medical letter can be converted, cut down to the paragraphs you need, and only those shared.

The honest part: where it stops

A scanned PDF gives an empty file. Its pages are pictures, and there is no text to find: our scanned test page produced a Markdown file of 1 byte and the warning “has no text at all: probably scanned. Run OCR PDF first to make the text readable, then convert it.” Do exactly that: OCR PDF, then convert the result. The guide to making scans searchable covers the OCR step.

Tables are found from how the text lines up, so check each one. A table with cells in clear columns, or drawn as a ruled grid, comes out as a Markdown table; a frame drawn around a block of plain text does not count as one. For a table of numbers you will work with, PDF to Excel keeps amounts as numbers. PDF to Word keeps tables too, if Word is where the document is going.

Heavily designed pages (brochures, magazines, slide decks saved as PDF) put text in boxes all over the page, and the result reads as a heap of fragments. Bold and italic did not survive in our test, and links and pictures do not either. So convert 1 file, read it next to the PDF, and only then convert the other 49. For the wider choice between Word, Excel and text, see PDF to Word, Excel or text.

Questions

Are tables kept as Markdown tables?

Yes, when the converter finds the table: cells in clear columns, or a ruled grid. In our test a 4-row table came out as a Markdown table with a header row. A frame drawn around plain text is not a table. Check each table against the PDF; for numbers you will calculate with, PDF to Excel is the better route.

Why are the page numbers and headers gone?

Lines that repeat at the top or bottom of most pages, such as running headers, footers and page numbers, are left out by default. Set Running lines to kept if you want them.

Why is my Markdown file empty?

The PDF is probably a scan: its pages are pictures with no text in them. The tool says so, for example has no text at all: probably scanned. Run OCR PDF first, then convert the result.

Can I tell which page a passage came from?

Turn Page markers on. Each page then starts with a comment such as page 2 in HTML comment form, which notes apps and Markdown viewers do not show but which you can search for.

Is my PDF uploaded to convert it?

No. The converter runs in your browser and the PDF stays on your device. If you then paste the Markdown into an online assistant or import it into Notion, that text goes to them, and you choose what goes.

The tools for this job