Extract Tables from PDFs in Python as Markdown
npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →Extract and validate the document first, then choose a parser for the table representation. GFM supports rectangular rows; complex HTML tables require a parser that understands row and column spans.
Convert a small folder first
The Python batch recipe saves one complete Markdown file per selected PDF, with durable input snapshots and identities for explicit recovery. Apply the local parser below to an individual output. For invoice line items, follow the invoice workflow before treating table cells as accounting data.
Select a PDF extraction method
Local table libraries expose different controls; they are not universally unable to handle whitespace or merged headers. pdfplumber supports line, text and explicit strategies. Camelot’s default rendering backend is pdfium, with Ghostscript optional. See the current primary-source comparison, checked 10 September 2026.
Inspect the dated source PDF and complete API output, then use the local parsing recipe below.
Parsing simple GFM output into pandas
First save validated conversion output as document.md. Install pandas in your Python environment, then save this script as table_to_csv.py and run it with Python 3. It reads a local file and exports only the first valid table candidate as CSV to stdout. If that first candidate is complex HTML, or no table exists, it fails explicitly. Later tables are not exported.
from pathlib import Path
import pandas as pd
# BEGIN SYNCED GFM TABLE PARSER
from html import unescape
def split_gfm_row(line):
row = line.strip()
if not (row.startswith("|") and row.endswith("|")):
raise ValueError("Expected a GFM table row with outer pipes")
cells, cell = [], []
escaped = False
for char in row[1:-1]:
if escaped:
if char in ("\\", "|", "*", "_", "`"):
cell.append(char)
else:
cell.extend(("\\", char))
escaped = False
elif char == "\\":
escaped = True
elif char == "|":
cells.append(unescape("".join(cell).strip()))
cell = []
else:
cell.append(char)
if escaped:
cell.append("\\")
cells.append(unescape("".join(cell).strip()))
return cells
def parse_gfm_candidate(lines):
if len(lines) < 2:
return None
rows = [split_gfm_row(line) for line in lines]
width = len(rows[0])
if width == 0 or len(rows[1]) != width:
return None
if any(cell != "---" for cell in rows[1]):
return None
if any(len(row) != width for row in rows[2:]):
return None
return rows
def iter_table_candidates(markdown):
current = []
for line in markdown.splitlines():
row = line.strip()
if row.startswith("|") and row.endswith("|"):
current.append(row)
continue
if current:
candidate = parse_gfm_candidate(current)
if candidate is not None:
yield "gfm", candidate
current = []
if "<table" in row.casefold():
yield "html", markdown
if current: # flush a table at end of output
candidate = parse_gfm_candidate(current)
if candidate is not None:
yield "gfm", candidate
def extract_gfm_tables(markdown):
return [value for kind, value in iter_table_candidates(markdown) if kind == "gfm"]
def first_table_branch(markdown):
return next(iter_table_candidates(markdown), ("none", None))
# END SYNCED GFM TABLE PARSER
markdown = Path("document.md").read_text(encoding="utf-8")
table_kind, table_rows = first_table_branch(markdown)
if table_kind == "html":
raise ValueError("Complex sanitized HTML table found: use a non-rendering HTML parser; do not render it as trusted HTML.")
if table_kind == "none":
raise ValueError("No valid table found")
header = table_rows[0]
data = table_rows[2:]
df = pd.DataFrame(data, columns=header)
print(df.to_csv(index=False), end="")
The helper splits structural pipes before unescaping cells. It requires a header, canonical separator and consistent row widths. Treat cell text as data. For complex sanitized HTML, use a non-rendering HTML parser that retains span attributes; do not render API output as trusted HTML.
Side-by-side summary
| Camelot | tabula-py | pdfplumber | pdfToMarkdown | |
|---|---|---|---|---|
| Output format | pandas DataFrame | pandas DataFrame | List of lists | Escaped GFM or sanitized HTML |
Choose based on your documents, processing requirements and downstream format. Keep CSV cells as strings until your application validates dates, numeric formats and currencies.
Convert a document
Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.
The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.
Validated Python conversion
Python 3 with requests. Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY. Allow up to 11 minutes; stopping the client can leave processing and charging running.
import base64
import os
import sys
from pathlib import Path
import requests
key = os.environ.get("PDFTOMARKDOWN_API_KEY")
identity = os.environ.get("PDFTOMARKDOWN_IDEMPOTENCY_KEY")
if not key or not identity:
sys.exit("Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY")
try:
response = requests.post(
"https://pdftomarkdown.dev/v1/convert",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": identity,
"Content-Type": "application/pdf"},
data=Path("document.pdf").read_bytes(),
timeout=(10, 660),
)
if response.status_code != 200:
sys.exit(f"HTTP {response.status_code}; keep the identity for recovery")
result = response.json()
if not (isinstance(result, dict) and result.get("complete") is True
and isinstance(result.get("markdown"), str)
and type(result.get("pages")) is int and result["pages"] >= 0
and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
except (requests.RequestException, ValueError, OSError):
sys.exit("Conversion failed; keep the identity and original input for recovery")