Learn / updated 2026-10-04
How PDF to Markdown conversion works: text layer, layout and OCR
A PDF is not a document in the way a Word file is. It is a set of drawing instructions: "put this text at this position, draw this line here". Nothing in the file says "this is a heading" or "these five lines are a table". Converting a PDF to Markdown means rebuilding that structure from positions.
First question: does the PDF have a text layer?
- Born-digital PDFs (exported from a word processor, an accounting system or a website) contain the text as characters with positions. Extraction is exact: the characters are right, and the work is rebuilding structure.
- Scanned PDFs are pictures of pages. There is no text to extract; a tool has to read the picture (OCR), and OCR can make mistakes with numbers, small print and poor scans.
A quick test: try to select text in a PDF viewer. If you can, it has a text layer.
Three approaches
- Layout rules. Group characters into lines by their vertical position, split lines into cells at wide gaps, treat repeated column positions as tables, and use font size to find headings. It is deterministic (the same file gives the same result), fast, cheap and easy to explain. It struggles with unusual layouts such as merged cells or multi-column pages.
- OCR. Needed for scans. A recognition engine turns pixels into text, which can then go through the same layout rules. Always check numbers.
- AI models. A language or vision model reads the page and writes the structure. It can cope with odd layouts, but results can differ between runs, errors can look plausible, and cost grows with the length of the output.
For invoices, reports and other documents you process in volume, layout rules plus OCR only where needed is often the better balance: predictable and cheap, with the model reserved for the cases rules cannot handle.
What good output looks like
- headings, lists and paragraphs in reading order,
- tables with the right number of rows and columns, including cells whose text wraps onto a second line and number columns that are right-aligned,
- fields (invoice number, dates, VAT, total) tied to the line they came from,
- a clear warning where OCR was used or a field was not found.
Check before you rely on it
Compare totals with the document, look at any page flagged as scanned, and treat a field filled from layout rather than a label (for example the issuer taken from the first line) as a suggestion.
Try it
The free PDF to Markdown tool runs in your browser and the file is never uploaded; the DocToJSON Actor does the same for batches and reads scans with OCR. For the reasoning behind rule-based extraction, see PDF and DOCX to Markdown without an LLM.
Frequently asked questions
Why do some PDFs convert badly? Layout, not content: columns, merged cells and decorative text positions are ambiguous without the original document structure.
Is Markdown the right target? For text, headings and lists, yes. Tables are best kept as rows or CSV, which is why good tools return both.
Related
- CPV codes and OCDS explained: how public tenders are described
- How EU VAT number validation works (VIES explained)
- SPF, DKIM and DMARC explained: how e-mail authentication works
- What is an LEI (Legal Entity Identifier)? Format, use and lookup
- All tools and prices
Informational only, not financial, legal or tax advice.