PDF to Markdown

Read a PDF from your terminal, app, or coding agent. Try page 1 without an account.

terminal
npx pdftomarkdown@0.1.4 https://pdftomarkdown.dev/samples/invoice.pdf

Node 18+. Published CLI 0.1.4 prints Markdown to your terminal. Saving files and account requests

Captured invoice output

SALESPERSON P.O. NUMBER REQUISITIONER SHIPPED VIA F.O.B. POINT TERMS
PO-3333
QUANTITYDESCRIPTIONUNIT PRICETOTAL
1Test for 23 fields1$100.00
SUBTOTAL$100.00
SALES TAX$10.00
TOTAL$110.00
PREVIOUS BALANCE$500.00
TOTAL DUE$610.00

Inspect the source and complete result

API JSON response
{
  "complete": true,
  "markdown": "# CONTOSO LTD.\n\n# INVOICE\n\nContoso Headquarters\n\n123 456th St\n\nNew York, NY, 10001\n\nINVOICE: INV-100\n\nDATE: 11/15/2019\n\nDUE DATE: 12/15/2019\n\nCUSTOMER NAME: MICROSOFT CORPORATION\n\nCUSTOMER ID: CID-12345\n\nMicrosoft Corp\n\n123 Other St,\n\nRedmond WA, 98052\n\nBILL TO:\n\nMicrosoft Finance\n\n123 Bill St,\n\nRedmond WA, 98052\n\nSHIP TO:\n\nMicrosoft Delivery\n\n123 Ship St,\n\nRedmond WA, 98052\n\nSERVICE ADDRESS:\n\nMicrosoft Services\n\n123 Service St,\n\nRedmond WA, 98052\n\n| SALESPERSON | P.O. NUMBER | REQUISITIONER | SHIPPED VIA | F.O.B. POINT | TERMS |\n| --- | --- | --- | --- | --- | --- |\n|  | PO-3333 |  |  |  |  |\n\n<table><tr><th>QUANTITY</th><th>DESCRIPTION</th><th>UNIT PRICE</th><th>TOTAL</th></tr><tr><td>1</td><td>Test for 23 fields</td><td>1</td><td>$100.00</td></tr><tr><td></td><td></td><td></td><td></td></tr><tr><td colspan=\"3\">SUBTOTAL</td><td>$100.00</td></tr><tr><td colspan=\"3\">SALES TAX</td><td>$10.00</td></tr><tr><td colspan=\"3\">TOTAL</td><td>$110.00</td></tr><tr><td colspan=\"3\">PREVIOUS BALANCE</td><td>$500.00</td></tr><tr><td colspan=\"3\">TOTAL DUE</td><td>$610.00</td></tr></table>\n\nTHANK YOU FOR YOUR BUSINESS!\n\nREMIT TO:\n\nContoso Billing\n\n123 Remit St\n\nNew York, NY, 10001\n\n> Processed by pdfToMarkdown.dev",
  "pages": 1,
  "request_id": "req_example_invoice"
}

What it handles

Convert document layouts, including Japanese and Chinese text, into Markdown. See the source PDF and complete invoice output, including purchase-order fields and merged table cells.

The API currently uses Mistral OCR 4.1 and returns a complete Markdown response. The PaddleOCR provider page is retained as a historical implementation guide.

Invoices & receipts

Research papers

Legal contracts

Financial reports

Japanese & Chinese docs

Scanned documents

Locales日本語ページ中文页面

Pricing

Try page 1 free, or get 20 trial pages once with a new account. Then choose paid credits. Monthly pages expire; top-ups do not.

Frequently asked questions

Is the PDF to Markdown API free?
Yes — the page-one preview and new-account trial are free. The Hacker tier needs no signup at all: the public key demo_public_key converts page 1 of any PDF at 3 requests per minute. The Developer tier (GitHub login, no credit card) gives a personal key with 20 trial pages once for new accounts and full multi-page conversion.
How do I convert a PDF to markdown without signing up?
Send one HTTP request with the public demo key, or run npx pdftomarkdown@0.1.4 document.pdf in a terminal (Node 18+, nothing to install). Both work immediately — the demo tier processes page 1 of the document.
Does it work on scanned PDFs?
Yes. A vision-language model performs real OCR on every page, so image-only PDFs convert the same way as digital ones — including tables and multi-column layouts. Explore the evaluation methodology.
How accurate is it on Japanese and Chinese documents?
Mistral OCR 4.1 extracts Japanese and Chinese text, headings and tables. Inspect a complete conversion or explore the evaluation methodology.
How long does a conversion take?
Processing time varies with document length, layout, and provider capacity. The call is synchronous and returns finished Markdown; set your HTTP client timeout to at least 11 minutes for large documents, or bound work with input.max_pages.
Can coding agents like Claude Code, Codex, or Cursor use it?
Yes — that's a primary use case. Claude Code has an official plugin; Codex and Cursor read PDFs via the CLI with a copy-paste AGENTS.md block or Cursor rule. Agents can also discover the API themselves through llms.txt and the OpenAPI spec.