Unstructured Alternative for RAG: When to Swap partition_pdf for an API
npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →Choose Unstructured when your pipeline needs typed elements or local processing. Choose pdfToMarkdown when it needs a complete Markdown response. Unstructured also offers hosted processing, so avoiding a local OCR stack alone does not decide between them.
Understand partition_pdf
Unstructured partitions documents into typed elements, including titles, narrative text and tables. Its PDF strategies include auto, fast, hi_res and ocr_only; the right strategy depends on the source and installed dependencies. High-resolution processing supports table extraction. Default behavior is not evidence that OCR is absent. Partitioning, strategies, checked 10 September 2026.
If a table is returned as HTML metadata, retain that structure when converting it for a downstream consumer. Filtering by element category can discard useful content, so inspect a sample before excluding entire categories.
Local control and hosted processing
A local Unstructured deployment lets you operate the document-processing environment. Dependencies and throughput depend on the selected strategy and hardware; this page does not assign a universal installation time or image size.
Unstructured also offers hosted processing through its API. Its documentation now recommends pipeline jobs for production use and labels the older Partition Endpoint as a legacy prototyping interface. Hosted processing options, checked 10 September 2026.
pdfToMarkdown processes PDFs externally and returns a complete Markdown string. Simple rectangular tables become escaped GFM; spans and nested table structures remain sanitized HTML. Review retention before sending documents.
Choose by what the next step consumes
Suppose you are adding an invoice to a retrieval pipeline. Start with the public source PDF and its complete, dated Markdown output. The pair lets you inspect the actual headings, line items and totals returned by pdfToMarkdown; it is not a paired Unstructured benchmark.
If your next step routes titles, narrative text and tables to different handlers, Unstructured’s typed elements give you those categories to work with. Keep table HTML metadata where available. You still need to decide how to serialize elements and carry their metadata into chunks; simply joining their text can lose structure. Local processing gives you control of the runtime, while hosted processing moves that operation to the provider.
If your next step stores a readable document or chunks Markdown headings, pdfToMarkdown supplies the complete string directly. You must still handle HTML tables, preserve a source identifier and validate retrieval on your corpus. The response is document content, not a validated invoice-field schema or a page-to-character citation map. Processing is external, so confirm that your document policy permits it.
Compare your actual workload
Test reading order, missing cells and the amount of application-specific processing. Our July 2026 CJK benchmark covers the earlier PaddleOCR processor and Tesseract, not an end-to-end Unstructured comparison or the current Mistral processor.
Use the current catalog for pdfToMarkdown costs. Measure runtime on your own documents.
Convert a document
Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.
The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.