· Updated · pdfToMarkdown team

Extract Tables from PDFs in Python as Markdown

pythontablesguidespdf
Try the API without signing up:npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →

Extract and validate the document first, then choose a parser for the table representation. GFM supports rectangular rows; complex HTML tables require a parser that understands row and column spans.

Convert a small folder first

The Python batch recipe saves one complete Markdown file per selected PDF, with durable input snapshots and identities for explicit recovery. Apply the local parser below to an individual output. For invoice line items, follow the invoice workflow before treating table cells as accounting data.

Select a PDF extraction method

Local table libraries expose different controls; they are not universally unable to handle whitespace or merged headers. pdfplumber supports line, text and explicit strategies. Camelot’s default rendering backend is pdfium, with Ghostscript optional. See the current primary-source comparison, checked 10 September 2026.

Inspect the dated source PDF and complete API output, then use the local parsing recipe below.

Parsing simple GFM output into pandas

First save validated conversion output as document.md. Install pandas in your Python environment, then save this script as table_to_csv.py and run it with Python 3. It reads a local file and exports only the first valid table candidate as CSV to stdout. If that first candidate is complex HTML, or no table exists, it fails explicitly. Later tables are not exported.

from pathlib import Path
import pandas as pd

# BEGIN SYNCED GFM TABLE PARSER
from html import unescape

def split_gfm_row(line):
    row = line.strip()
    if not (row.startswith("|") and row.endswith("|")):
        raise ValueError("Expected a GFM table row with outer pipes")

    cells, cell = [], []
    escaped = False
    for char in row[1:-1]:
        if escaped:
            if char in ("\\", "|", "*", "_", "`"):
                cell.append(char)
            else:
                cell.extend(("\\", char))
            escaped = False
        elif char == "\\":
            escaped = True
        elif char == "|":
            cells.append(unescape("".join(cell).strip()))
            cell = []
        else:
            cell.append(char)
    if escaped:
        cell.append("\\")
    cells.append(unescape("".join(cell).strip()))
    return cells

def parse_gfm_candidate(lines):
    if len(lines) < 2:
        return None
    rows = [split_gfm_row(line) for line in lines]
    width = len(rows[0])
    if width == 0 or len(rows[1]) != width:
        return None
    if any(cell != "---" for cell in rows[1]):
        return None
    if any(len(row) != width for row in rows[2:]):
        return None
    return rows

def iter_table_candidates(markdown):
    current = []
    for line in markdown.splitlines():
        row = line.strip()
        if row.startswith("|") and row.endswith("|"):
            current.append(row)
            continue
        if current:
            candidate = parse_gfm_candidate(current)
            if candidate is not None:
                yield "gfm", candidate
            current = []
        if "<table" in row.casefold():
            yield "html", markdown
    if current:  # flush a table at end of output
        candidate = parse_gfm_candidate(current)
        if candidate is not None:
            yield "gfm", candidate

def extract_gfm_tables(markdown):
    return [value for kind, value in iter_table_candidates(markdown) if kind == "gfm"]

def first_table_branch(markdown):
    return next(iter_table_candidates(markdown), ("none", None))
# END SYNCED GFM TABLE PARSER

markdown = Path("document.md").read_text(encoding="utf-8")
table_kind, table_rows = first_table_branch(markdown)
if table_kind == "html":
    raise ValueError("Complex sanitized HTML table found: use a non-rendering HTML parser; do not render it as trusted HTML.")
if table_kind == "none":
    raise ValueError("No valid table found")
header = table_rows[0]
data = table_rows[2:]
df = pd.DataFrame(data, columns=header)
print(df.to_csv(index=False), end="")

The helper splits structural pipes before unescaping cells. It requires a header, canonical separator and consistent row widths. Treat cell text as data. For complex sanitized HTML, use a non-rendering HTML parser that retains span attributes; do not render API output as trusted HTML.

Side-by-side summary

Camelot tabula-py pdfplumber pdfToMarkdown
Output format pandas DataFrame pandas DataFrame List of lists Escaped GFM or sanitized HTML

Choose based on your documents, processing requirements and downstream format. Keep CSV cells as strings until your application validates dates, numeric formats and currencies.

Convert a document

Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.

The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.

Validated Python conversion

Python 3 with requests. Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY. Allow up to 11 minutes; stopping the client can leave processing and charging running.

pythonRaw
import base64
import os
import sys
from pathlib import Path
import requests

key = os.environ.get("PDFTOMARKDOWN_API_KEY")
identity = os.environ.get("PDFTOMARKDOWN_IDEMPOTENCY_KEY")
if not key or not identity:
    sys.exit("Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY")
try:
    response = requests.post(
        "https://pdftomarkdown.dev/v1/convert",
        headers={"Authorization": f"Bearer {key}", "Idempotency-Key": identity,
                 "Content-Type": "application/pdf"},
        data=Path("document.pdf").read_bytes(),
        timeout=(10, 660),
    )
    if response.status_code != 200:
        sys.exit(f"HTTP {response.status_code}; keep the identity for recovery")
    result = response.json()
    if not (isinstance(result, dict) and result.get("complete") is True
            and isinstance(result.get("markdown"), str)
            and type(result.get("pages")) is int and result["pages"] >= 0
            and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
        sys.exit("Incomplete or invalid conversion response")
    sys.stdout.write(result["markdown"])
except (requests.RequestException, ValueError, OSError):
    sys.exit("Conversion failed; keep the identity and original input for recovery")