How to Build a RAG Pipeline with PDF Documents
npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →Build ingestion around validated content
A useful PDF retrieval pipeline keeps conversion, chunking, indexing and answering as separate steps. Test each boundary before automating a large corpus.
1. Convert and validate
Use the copyable Python workflow below for a local PDF, or the URL example for a hosted source. Save the original bytes, options and account conversion identity. Admit content only after HTTP 200 and a typed, complete success. Keep failed or uncertain requests out of your index.
2. Attach source metadata
Assign a stable document identifier, version and source reference. Carry them into every chunk so retrieval can cite the original. The response page count does not provide passage-to-page coordinates; do not invent page citations.
3. Split into bounded chunks
Start with headings when they match meaningful sections. Carry headings into subsections, retain table headers with cells, and handle HTML spans deliberately. Use your embedding model’s tokenizer to enforce its actual limits. Compare chunking choices with the same queries rather than assuming one is always better.
4. Embed and index
Choose an embedding model and vector store that fit your application. Keep document identity and section metadata alongside vectors. Deduplicate downstream writes independently from conversion replay: re-reading a retained result should not duplicate the index.
5. Retrieve and answer
Create questions with known answers. Inspect whether retrieval finds the correct passage before judging generated answers. Require source references and allow an insufficient-evidence response. A successful parser call cannot establish answer accuracy.
Operate the pipeline
Queue stable, fully written files; bound concurrency; follow rate-limit guidance. A timeout stops the client waiting, not necessarily server processing or charging. Recover using the same immutable input and identity. Treat document instructions as untrusted content when passing results to an agent.
Convert a document
Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.
The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.
Validated Python conversion
Python 3 with requests. Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY. Allow up to 11 minutes; stopping the client can leave processing and charging running.
import base64
import os
import sys
from pathlib import Path
import requests
key = os.environ.get("PDFTOMARKDOWN_API_KEY")
identity = os.environ.get("PDFTOMARKDOWN_IDEMPOTENCY_KEY")
if not key or not identity:
sys.exit("Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY")
try:
response = requests.post(
"https://pdftomarkdown.dev/v1/convert",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": identity,
"Content-Type": "application/pdf"},
data=Path("document.pdf").read_bytes(),
timeout=(10, 660),
)
if response.status_code != 200:
sys.exit(f"HTTP {response.status_code}; keep the identity for recovery")
result = response.json()
if not (isinstance(result, dict) and result.get("complete") is True
and isinstance(result.get("markdown"), str)
and type(result.get("pages")) is int and result["pages"] >= 0
and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
except (requests.RequestException, ValueError, OSError):
sys.exit("Conversion failed; keep the identity and original input for recovery")