ambolt

Blog / 2026-10-03

PDF and DOCX to Markdown and tables, without an LLM

Language models can read documents, but they give a different answer on a different day and charge by the token. For invoices, reports and contracts that you process in volume, a deterministic parser is often the better tool: the same file gives the same output, and the price does not depend on how long the answer is.

DocToJSON reads PDF and DOCX files from a public URL and returns three things.

Markdown with structure

Headings and lists become Markdown. Text keeps the order it has on the page.

Tables as data

Every table is returned as rows and as CSV, so it can go straight into a sheet or a database.

Named fields

You can pass hints such as invoice number, date or total. The tool finds a value by its label and returns it together with the line it came from. A field that is not found is reported as not found, never guessed.

Scans

Scanned pages are read with OCR (Tesseract). The output warns you on every page where OCR was used, because text from a scan deserves a second look.

Honest limits

Price

It is billed per page, with a higher price for OCR pages, and failed files are not charged. Run it as an Apify Actor; see the DocToJSON page for inputs, outputs and current prices.

More from the blog

Data and prices change; every API response states its source and date. Informational only, not financial, legal or tax advice.