services / doc-parse
textoperational · 614 ms
Document to Markdown (parse)
Turn a document into clean, agent-ready Markdown with metadata: DOCX, ODT, EPUB, RTF, HTML, RST, LaTeX and more. Returns title, author, date, language, word count, heading outline, table and image counts and the Markdown (headings and tables preserved). For PDF use pdf-extract; PPTX and XLSX are not supported. Params: file_base64 or url or content (one required), from (optional), max_chars (1000 to 200000, default 50000).
Run free trial ↗1 free call per day with the example input. Paid: $0.006 USDC, no limit.
Call it
Input
| Field | Type | Description |
|---|---|---|
| file_base64 | string | Document as base64 (max 10 MB) |
| url | string | |
| content | string | |
| from | string | |
| max_chars | integer = 50000 |
Output
| Field | Type | Description |
|---|---|---|
| format | string | |
| source | string | |
| title | string | |
| author | any | |
| date | any | |
| language | string | |
| word_count | integer | |
| characters | integer | |
| heading_count | integer | |
| table_count | integer | |
| image_count | integer | |
| outline | array | |
| markdown | string | |
| truncated | boolean |
Example response (data)
{
"format": "html",
"source": "https://agentsvc.io/docs",
"title": "Docs: how agents pay per call with x402 · agentsvc.io",
"author": null,
"date": null,
"language": "en",
"word_count": 8068,
"characters": 2000,
"heading_count": 8,
"table_count": 0,
"image_count": 0,
"outline": [
{
"level": 1,
"text": "Pay per call. Nothing else to set up."
},
{
"level": 2,
"text": "Free trial, no wallet"
},
{
"level": 1,
"text": "GET works too and uses the example input:"
}
],
"markdown": "<div hidden=\"\">\n\n</div>\n\n<div style=\"border-bottom:1px solid var(--border);background-color:rgba(10, 14, 26, 0.92);backdrop-filter:blur(12px);position:sticky;top:0;z-index:50\">\n\n<div class=\"container\" style=\"display:flex;align-items:center;justify-content:space-between;min-height:60px;gap:1rem\">\n…",
"truncated": true
}