URL to Markdown: Clean Page Text for LLMs and RAG
Fetch a public web page and return its main text as Markdown with title, headings, lists, tables and absolute links. Respects robots.txt, refuses private and internal addresses, no JavaScript rendering. For pages that need a browser, use a rendering tool.
x402: $0.003 per callMCP tool: url_to_markdown
Input
| Parameter | Type | Default | Description |
|---|---|---|---|
url * | string | Public http(s) URL. | |
maxChars | integer | 20000 | Maximum characters of Markdown returned. |
* required
Source and freshness
- Source and licence
- Stated in every response (field source).
- Last verified
- 2026-10-03 (hourly check against the live source)
Use it
HTTP (x402)
curl -i "https://api.ambolt.dev/v1/url-to-markdown?url=https%3A%2F%2Fwww.sitemaps.org%2F"Returns 402 Payment Required with the price until an x402 payment is attached; @x402/fetch does this for you. See Get started.
MCP
{ "mcpServers": { "ambolt": { "url": "https://api.ambolt.dev/mcp" } } }
# then call the tool: url_to_markdown
Apify
Not published as an Actor yet.
Example response
{
"url": "https://www.sitemaps.org/",
"title": "sitemaps.org - Home",
"markdown": "# sitemaps.org\n\n\r\n\r\n\n\r\n\r\n\r\n\r\n\r\n\r\n\n- [FAQ](https://www.sitemaps.org/faq.php)\r\n\r\n\n- [Protocol](https://www.sitemaps.org/protocol.php)\r\n\r\n\n- [Home](https://www.sitemaps.org/#)\r\n\r\n\n\r\n\r\n\n\r\n\r\n\r\n\r\n\n\r\n\r\n\r\n\r\n\r\n\r\nLanguage: \r\n\r\n\n\r\n\r\n\r\n\r\n\n# What are Sitemaps?\n\n\r\n\r\n\r\n\r\nSitemaps are an easy way for webmasters to inform search engines about pages on\r\n\r\ntheir sites that are available for crawling. In its simplest form, a Sitemap is\r\n\r\nan XML file that lists URLs for a site along with additional metadata about each\r\n\r\nURL (when it was last updated, how often it usually changes, and how important it\r\n\r\nis, relative to other URLs in the site) so that search engines can more intelligently\r\n\r\ncrawl the site.\n\n\r\n\r\n\r\n\r\nWeb crawlers usually discover pages from links within the site and from other sites.\r\n\r\nSitemaps supplement this data to allow crawlers that support Sitemaps to pick up\r\n\r\nall URLs in the Sitemap and learn about those URLs using the associated metadata.\r\n\r\nUsing the Sitemap [protocol](https://www.sitemaps.org/protocol.php) does not guarantee that web\r\n\r\npages are included in search engines, but provides hints for web crawlers to do\r\n\r\na better job of crawling your site.\n\n\r\n\r\n\r\n\r\nSitemap 0.90 is offered under the terms of the [Attribution-ShareAlike Creative Commons License](http://creativecommons.org/licenses/by-sa/2.5/) and has wide adoption, including\r\n\r\nsupport from Google, Yahoo!, and Microsoft.\n\n\r\n\r\n\r\n\r\nLast Updated: 17 April 2020\r\n\r\n\n\r\n\r\n\n\r\n\r\n\r\n\r\n\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n[Terms and conditions](https://www.sitemaps.org/terms.php)",
"characters": 1544,
"truncated": false,
"fetchedAt": "2026-10-02T23:25:07.738Z",
"note": "Extracted text is the publisher's content; check its licence before reuse."
}
Good to know
- You are charged only when the call succeeds. Invalid input or an unavailable source costs nothing.
- Every response states its source and the time it was fetched.
- Informational only; not financial, legal or tax advice.
- Something wrong or missing? Open an issue on the repository (ambolt-mcp) and a reply follows under the Ambolt name.