Document AI Without Fine-Tuning: How Vision-Language Models Changed OCR
npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →Historical architecture note. The earlier pdfToMarkdown integration used PaddleOCR. The current processor is Mistral OCR 4.1. The historical provider guide retains implementation and model references.
Why visual document processing helps
Rendered pages contain layout cues: columns, spacing, typography, table boundaries and captions. A document model can use these cues to produce text with structure, including when the PDF lacks a usable text layer.
That does not establish that every model understands every layout. Small print, unusual notation and degraded scans can still cause transcription or ordering errors. A model update requires evaluation on representative documents; replacing a dependency alone does not establish better output.
Keep transcription and field extraction separate
The conversion API returns document content as Markdown. Mapping it into an invoice, clinical or contract schema remains an application responsibility. Define required fields, preserve unknowns and validate extracted values against the source.
Evaluate delivery and quality separately
A canonical success marks a complete response, not perfect OCR. Check both: reject malformed or incomplete API payloads, then evaluate content accuracy and structure on your corpus. The dated invoice capture supplies a source/output pair; the July 2026 CJK benchmark remains historical PaddleOCR evidence.
Convert a document
Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.
The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.