Python Invoice OCR: Extract PDF Invoices to Markdown and CSV
npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →Separate invoice OCR from accounting writes
Convert and validate invoice content before extracting fields. A conversion result is not an approved bill, and a syntactically valid JSON extraction is not verified accounting data.
1. Ingest a stable file
Wait until an uploaded or watched file is complete before enqueueing it. Store its identity and conversion options, then save an idempotency key before account submission. Duplicate filesystem events must not create duplicate conversion jobs.
2. Convert with the tested Python recipe
The bounded Python batch recipe reads up to five selected local PDFs, snapshots their bytes and options, and saves an identity before each submission. It validates HTTP status, complete Markdown, page count and request ID, then writes each output without overwriting an existing file. For timeout or transport uncertainty, retain the original bytes and identity; follow the canonical retry guidance before starting new work.
3. Extract a defined schema
Define vendor, invoice number, currency, dates, line items and totals explicitly. Use deterministic rules where labels are stable. If using an LLM for variable layouts, validate its output against the same schema and retain missing fields as unknown. No current evidence supports a universal regex hit rate or a claim that hallucinations disappear.
4. Reconcile amounts
Use decimal arithmetic for money. Check line-item totals, subtotal, discounts, tax and final amount. Keep credit notes and negative amounts distinct. Route mismatches and ambiguous currency or dates to review.
5. Export reviewed records
Use an accounting provider’s current schema and sandbox to test record creation. Keep business-level duplicate detection separate from API replay. Review CSV content before opening it in spreadsheet software, since document text can contain formulas.
What this guide verifies
The conversion recipe runs against local success and failure fixtures. It does not claim a tested QuickBooks, Xero, folder-watcher or LLM integration. Those require application-specific schemas, permissions and tests. The invoice capture supplies a real source/output pair for inspection.
Convert a document
Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.
The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.
Validated Python conversion
Python 3 with requests. Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY. Allow up to 11 minutes; stopping the client can leave processing and charging running.
import base64
import os
import sys
from pathlib import Path
import requests
key = os.environ.get("PDFTOMARKDOWN_API_KEY")
identity = os.environ.get("PDFTOMARKDOWN_IDEMPOTENCY_KEY")
if not key or not identity:
sys.exit("Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY")
try:
response = requests.post(
"https://pdftomarkdown.dev/v1/convert",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": identity,
"Content-Type": "application/pdf"},
data=Path("document.pdf").read_bytes(),
timeout=(10, 660),
)
if response.status_code != 200:
sys.exit(f"HTTP {response.status_code}; keep the identity for recovery")
result = response.json()
if not (isinstance(result, dict) and result.get("complete") is True
and isinstance(result.get("markdown"), str)
and type(result.get("pages")) is int and result["pages"] >= 0
and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
except (requests.RequestException, ValueError, OSError):
sys.exit("Conversion failed; keep the identity and original input for recovery")