PDF to Markdown API

Convert native and scanned PDFs to Markdown with one REST endpoint built for developers, RAG pipelines, and document automation.

Inspect the output first, then choose your input and page budget.

Try a PDF · Make your first API request · Endpoint reference

Response

{
  "complete": true,
  "markdown": "# CONTOSO LTD.\n\n# INVOICE\n\nContoso Headquarters\n\n123 456th St\n\nNew York, NY, 10001\n\nINVOICE: INV-100\n\nDATE: 11/15/2019\n\nDUE DATE: 12/15/2019\n\nCUSTOMER NAME: MICROSOFT CORPORATION\n\nCUSTOMER ID: CID-12345\n\nMicrosoft Corp\n\n123 Other St,\n\nRedmond WA, 98052\n\nBILL TO:\n\nMicrosoft Finance\n\n123 Bill St,\n\nRedmond WA, 98052\n\nSHIP TO:\n\nMicrosoft Delivery\n\n123 Ship St,\n\nRedmond WA, 98052\n\nSERVICE ADDRESS:\n\nMicrosoft Services\n\n123 Service St,\n\nRedmond WA, 98052\n\n| SALESPERSON | P.O. NUMBER | REQUISITIONER | SHIPPED VIA | F.O.B. POINT | TERMS |\n| --- | --- | --- | --- | --- | --- |\n|  | PO-3333 |  |  |  |  |\n\n<table><tr><th>QUANTITY</th><th>DESCRIPTION</th><th>UNIT PRICE</th><th>TOTAL</th></tr><tr><td>1</td><td>Test for 23 fields</td><td>1</td><td>$100.00</td></tr><tr><td></td><td></td><td></td><td></td></tr><tr><td colspan=\"3\">SUBTOTAL</td><td>$100.00</td></tr><tr><td colspan=\"3\">SALES TAX</td><td>$10.00</td></tr><tr><td colspan=\"3\">TOTAL</td><td>$110.00</td></tr><tr><td colspan=\"3\">PREVIOUS BALANCE</td><td>$500.00</td></tr><tr><td colspan=\"3\">TOTAL DUE</td><td>$610.00</td></tr></table>\n\nTHANK YOU FOR YOUR BUSINESS!\n\nREMIT TO:\n\nContoso Billing\n\n123 Remit St\n\nNew York, NY, 10001\n\n> Processed by pdfToMarkdown.dev",
  "pages": 1,
  "request_id": "req_example_invoice"
}

A success is explicitly marked complete: true. Simple rectangular tables use escaped GFM pipe-table syntax. Complex tables with spans or nested content remain sanitized CommonMark-compatible raw HTML so their structure is not discarded. Over-budget complete output fails with response_too_large instead of returning partial Markdown.

Compare the dated source PDF and full output before integrating. The API returns Markdown; your application separately extracts and validates any business fields.

Choose input, limits and cost

  • Use a public PDF URL, Base64 bytes or a raw PDF upload. Raw uploads avoid Base64 expansion.
  • All inputs share a 50 MiB decoded-file limit and a 1,000-page document limit. Set max_pages to bound the selected prefix.
  • The public demo converts page 1 with a watermark at 3 requests per minute per IP. New accounts receive 20 trial pages once.
  • Each delivered account page uses one credit. See current paid credits; monthly pages expire and independent top-ups do not.

Allow up to 11 minutes for the synchronous request. Processing time varies; a client timeout can leave processing and charging in progress. Save the input, options and identity before submitting account work; the Python batch recipe demonstrates recovery.

When to use it

Use the API when Markdown is the handoff format for chunking, embeddings, LLM prompts, search indexing, document review, or downstream extraction. If you need local-only processing, use a local parser instead.

The API currently uses Mistral OCR 4.1 with the same provider-neutral response contract. The PaddleOCR provider page documents the earlier hosted implementation.

Frequently asked questions

What are the rate limits and quotas?

The keyless demo tier allows 3 requests per minute per IP and converts page 1 only. A free Developer key (GitHub login) gives 20 trial pages once for new accounts with full multi-page support. Rate-limit responses include retry guidance; exhausted starter credits do not renew monthly.

Can I call the API from browser JavaScript?

Yes — CORS is enabled (Access-Control-Allow-Origin: *), and the timing and request-id headers are exposed to browser code.

How do I keep latency bounded on large PDFs?

Set input.max_pages to cap the pages processed, and give your HTTP client a timeout of 11 minutes. Timing headers are optional. Measure total elapsed time locally; provider latency is not total request time.

What do error responses look like?

Conversion errors include error (machine-readable code), message (states the recommended fix), and request_id; docs and retry fields appear where relevant. For example, an unreachable pdf_url returns 422 with a message suggesting the pdf_base64 fallback.

Is there an SDK?

No SDK is required — it's one HTTP call from any language, and the OpenAPI 3.1 spec can generate a typed client. There is also an official CLI (npx pdftomarkdown) and a Claude Code plugin.

Trust links

Review security, privacy, and data retention before processing sensitive documents.