Historical provider guide · PaddleOCR-VL-1.6

Historical PaddleOCR-VL API implementation

This page records the earlier managed GPU integration. It is not a claim about the provider selected for new requests.

What this is

pdfToMarkdown previously ran PaddleOCR-VL-1.6 on managed GPUs behind one REST endpoint. Send a PDF by public URL or base64 and receive finished Markdown plus page and request metadata. The stable keyless demo request remains current; the public tier processes page 1 and is limited to 3 requests per minute per IP.

curl -X POST https://pdftomarkdown.dev/v1/convert \
  -H "Authorization: Bearer demo_public_key" \
  -H "Content-Type: application/json" \
  -d '{"input":{"pdf_url":"https://pdftomarkdown.dev/samples/invoice.pdf"}}'

The model

PaddleOCR-VL-1.6 was released on May 28, 2026. Its model card describes a compact 0.9B vision-language model built on ERNIE-4.5-0.3B, with an Apache-2.0 license and an architecture compatible with PaddleOCR-VL-1.5. In thetechnical report, the model authors report 96.33% overall on OmniDocBench v1.6; that is an author-reported model result, not an end-to-end benchmark of this API.

Why use a hosted version

PaddleOCR-VL is open source, and official hosted alternatives exist. The practical difference is the account, infrastructure, billing, and output contract around the model. The access table below is a historical July 2026 snapshot; check the linked providers for current availability and prices.

OptionAccess and billingWhat you manage
Self-hostApache-2.0 model weightsGPU capacity, CUDA and FlashAttention runtime, deployment, monitoring, and version churn
Paddle AI StudioLogin and API token; 3,000 pages per model per day; 100 pages per fileAI Studio account, token, and service-specific limits
Baidu Intelligent CloudReal-name verification; RMB billing at ¥0.09 per pageBaidu Cloud account, credentials, and RMB billing
pdfToMarkdownKeyless demo → free key with GitHub login; global billing when paid tiers applyOne stable JSON/Markdown contract with documented error semantics

Output contract

  • Successful responses explicitly include complete: true.
  • Simple rectangular tables use escaped GFM pipe-table syntax where the structure is representable.
  • Tables with spans or nested content remain sanitized CommonMark-compatible HTML instead of losing cells.
  • Complete output is never silently truncated; over-budget responses fail with response_too_large.
  • Errors use a documented JSON shape with a machine-readable code, recommended fix, docs URL, and request ID.
  • See the API docs and OpenAPI 3.1 specification for the canonical schema.

Measured proof on Japanese and Chinese scans

In our reproducible CJK OCR benchmark, the API measured 0.10% character error rate on Japanese prose and 0.00% on Simplified Chinese prose, while preserving a 28-cell Japanese forecast table. Those API runs used PaddleOCR-VL-1.5; version 1.6 is architecture-compatible, but the benchmark remains labeled with the version actually tested. Read the inputs, methodology, raw excerpts, and limitations.

Frequently asked questions

Which model did this guide document?

This historical guide documents the PaddleOCR-VL-1.6 production integration. The current API keeps provider changes behind the same stable response contract.

Can I self-host instead?

Yes. PaddleOCR-VL-1.6 is available under Apache-2.0 from the official model card. Self-hosting gives you full infrastructure control; the historical hosted setup handled GPU capacity, the CUDA and FlashAttention runtime, model upgrades, PDF page processing, and the API contract.

Is PaddleOCR-VL-1.6 still a request option?

No per-request model selection is exposed. This page is retained as a historical record of the earlier hosted provider.

What are the limits?

Hacker: public demo key, page 1 only, 3 requests per minute per IP. Developer: free GitHub-login key, 20 trial pages once for new accounts, full multi-page support. See the current tier table in the docs.