· Updated · pdfToMarkdown team

Why Convert PDFs to Markdown for LLMs, RAG, and Search

guidespdfmarkdown
Try the API without signing up:npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →

Use Markdown when document structure helps the next step

Markdown provides readable headings, lists and tables for search, review and LLM workflows. Native PDF support varies by tool; conversion is useful when you need OCR or a reusable textual artifact.

A practical intermediate format

Store Markdown with its source identifier, inspect it directly, or parse it into a document tree. Keep complex tables as structured HTML when pipe rows would lose spans. Plain text can be sufficient for simple content; choose the format your workflow requires.

Make retrieval testable

Headings and tables can preserve context for chunks. Their presence alone does not prove better retrieval: compare known questions and source-backed answers on your own corpus.

Separate content from business data

The API returns document content. Extracting an invoice schema, identifying contract obligations or normalizing transactions requires additional validation. Inspect the dated invoice capture for a concrete output example.

Convert a document

Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.

The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.