Blog / 2026-10-03
PDF and DOCX to Markdown and tables, without an LLM
Language models can read documents, but they give a different answer on a different day and charge by the token. For invoices, reports and contracts that you process in volume, a deterministic parser is often the better tool: the same file gives the same output, and the price does not depend on how long the answer is.
DocToJSON reads PDF and DOCX files from a public URL and returns three things.
Markdown with structure
Headings and lists become Markdown. Text keeps the order it has on the page.
Tables as data
Every table is returned as rows and as CSV, so it can go straight into a sheet or a database.
Named fields
You can pass hints such as invoice number, date or total. The tool finds a value by its label and returns it together with the line it came from. A field that is not found is reported as not found, never guessed.
Scans
Scanned pages are read with OCR (Tesseract). The output warns you on every page where OCR was used, because text from a scan deserves a second look.
Honest limits
- Merged cells, multi-line headers and unusual layouts can need checking.
- Image-only pages give nothing unless OCR can read them.
- Files are fetched for the run and not kept beyond the platform's own run storage.
Price
It is billed per page, with a higher price for OCR pages, and failed files are not charged. Run it as an Apify Actor; see the DocToJSON page for inputs, outputs and current prices.
More from the blog
- EU and UK public tenders in one schema, with OCDS export
- How an agent pays per call with x402
- Resolve a company to its official record with one call
- All tools and prices
- Guides
Data and prices change; every API response states its source and date. Informational only, not financial, legal or tax advice.