Catalog / DocToJSON
PDF and DOCX files in; Markdown, CSV tables and named fields out.
Turn documents into structured data without a subscription. Headings and lists become Markdown, every table is also returned as rows and CSV, and fields such as invoice number, date and total are found by label. Scanned pages are read with OCR.
Input
- Public URL of a PDF or DOCX file
- Optional field hints (name, type, labels)
- Optional page limit
Output
- Markdown with headings, lists and tables
- Tables as rows and CSV
- Named fields with the line they came from
- A warning where OCR was used
Not returned
- Guessed values: a field that is not found is reported as not found
- Anything from image-only pages unless OCR can read them
Try it
apify call ambolt/doc-to-json -i '{"urls":["https://example.com/file.pdf"]}'Example response
{
"files": 1,
"extracted": 1,
"pagesProcessed": 4,
"billableResults": 4,
"datasetItems": [
{
"url": "https://www.rfc-editor.org/rfc/rfc9110.pdf",
"ok": true,
"kind": "pdf",
"pages": 4,
"characters": 5134,
"markdown": "| Stream: | Internet Engineering Task Force (IETF) |\n| --- | --- |\n| RFC: | 9110 |\n| STD: | 97 |\n| Obsoletes: | 2818, 7230, 7231, 7232, 7233, 7235, 7538, 7615, 7694 |\n| Updates: | 3864 |\n| Category: | Standards Track |\n| Published: | June 2022 |\n| ISSN: | 2070-1721 |\n| Authors: | R. Fielding, Ed. M. Nottingham, Ed. J. Reschke, Ed. |\n| | Adobe Fastly greenbytes |\n\n# RFC 9110\n\n# HTTP Semantics\n\n## Abstract\n\nThe Hypertext Transfer Protocol (HTTP) is a stateless application-level protocol for distributed, collaborative, hypertext information systems. This document describes the overall architecture of HTTP, establishes common terminology, and defines aspects of the protocol that are shared by all versions. In this definition are core protocol elements, extensibility mechanisms, and the \"http\" and \"https\" Uniform Resource Identifier (URI) schemes.\n\nThis document updates RFC 3864 and obsoletes RFCs 2818, 7231, 7232, 7233, 7235, 7538, 7615, 7694, and portions of 7230.\n\n## Status of This Memo\n\nThis is an Internet Standards Track document.\n\nThis document is a product of the Internet Engineering Task Force (IETF). It represents the consensus of the IETF community. It has received public review and has been approved for publication by the Internet Engineering Steering Group (IESG). Further information on Internet Standards is available in Section 2 of RFC 7841.\n\nInformation about the current status of this document, any errata, and how to provide feedback on it may be obtained at https://www.rfc-editor.org/info/rfc9110.\n\n## Copyright Notice\n\nCopyright (c) 2022 IETF Trust and the persons identified as the document authors. All rights reserved.\n\nFielding, et al. Standards Track Page 1\n\nRFC 9110 HTTP Semantics June 2022\n\nThis document is subject to BCP 78 and the IETF Trust's Legal \n…",
"tableCount": 1,
"tables": [
{
"page": 1,
"rowCount": 10,
"rows": [
[
"Stream:",
"Internet Engineering Task Force (IETF)"
],
[
"RFC:",
"9110"
…Recorded recently. Every real response names its source and the time it was read.
Pricing
| Rail | Unit | Price | Free allowance |
|---|---|---|---|
| Apify Actor | page | $0.03 per page; $0.06 per OCR page | Apify trial credits; failed files are not charged |
Failed calls are never charged. Prices are per successful call or event.
Questions
Is it AI?
No. It is rule-based layout analysis, so the same file gives the same result. OCR uses Tesseract.
What about complex layouts?
Merged cells, multi-line headers and scans can need checking; the output flags OCR pages.
Are my files stored?
Files are fetched for the run and not kept beyond Apify's own run storage.
Informational only, not financial, legal or tax advice. Something wrong or missing? Open an issue on the repository; a reply follows under the product name.