· Updated · pdfToMarkdown team

How to Build a RAG Pipeline with PDF Documents

ragtutoriallangchainpythonllm
Try the API without signing up:npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →

Build ingestion around validated content

A useful PDF retrieval pipeline keeps conversion, chunking, indexing and answering as separate steps. Test each boundary before automating a large corpus.

1. Convert and validate

Use the copyable Python workflow below for a local PDF, or the URL example for a hosted source. Save the original bytes, options and account conversion identity. Admit content only after HTTP 200 and a typed, complete success. Keep failed or uncertain requests out of your index.

2. Attach source metadata

Assign a stable document identifier, version and source reference. Carry them into every chunk so retrieval can cite the original. The response page count does not provide passage-to-page coordinates; do not invent page citations.

3. Split into bounded chunks

Start with headings when they match meaningful sections. Carry headings into subsections, retain table headers with cells, and handle HTML spans deliberately. Use your embedding model’s tokenizer to enforce its actual limits. Compare chunking choices with the same queries rather than assuming one is always better.

4. Embed and index

Choose an embedding model and vector store that fit your application. Keep document identity and section metadata alongside vectors. Deduplicate downstream writes independently from conversion replay: re-reading a retained result should not duplicate the index.

5. Retrieve and answer

Create questions with known answers. Inspect whether retrieval finds the correct passage before judging generated answers. Require source references and allow an insufficient-evidence response. A successful parser call cannot establish answer accuracy.

Operate the pipeline

Queue stable, fully written files; bound concurrency; follow rate-limit guidance. A timeout stops the client waiting, not necessarily server processing or charging. Recover using the same immutable input and identity. Treat document instructions as untrusted content when passing results to an agent.

Convert a document

Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.

The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.

Validated Python conversion

Python 3 with requests. Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY. Allow up to 11 minutes; stopping the client can leave processing and charging running.

pythonRaw
import base64
import os
import sys
from pathlib import Path
import requests

key = os.environ.get("PDFTOMARKDOWN_API_KEY")
identity = os.environ.get("PDFTOMARKDOWN_IDEMPOTENCY_KEY")
if not key or not identity:
    sys.exit("Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY")
try:
    response = requests.post(
        "https://pdftomarkdown.dev/v1/convert",
        headers={"Authorization": f"Bearer {key}", "Idempotency-Key": identity,
                 "Content-Type": "application/pdf"},
        data=Path("document.pdf").read_bytes(),
        timeout=(10, 660),
    )
    if response.status_code != 200:
        sys.exit(f"HTTP {response.status_code}; keep the identity for recovery")
    result = response.json()
    if not (isinstance(result, dict) and result.get("complete") is True
            and isinstance(result.get("markdown"), str)
            and type(result.get("pages")) is int and result["pages"] >= 0
            and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
        sys.exit("Incomplete or invalid conversion response")
    sys.stdout.write(result["markdown"])
except (requests.RequestException, ValueError, OSError):
    sys.exit("Conversion failed; keep the identity and original input for recovery")