pdf table extraction api

PDF Table Extraction API

One endpoint. POST a PDF, get complete Markdown back — simple tables as GFM, and complex table structure preserved as sanitized HTML when GFM would be lossy.

Server-side PDF processing. Review privacy, security, and data retention.

Convert selected files before parsing

Use the Python batch recipe to save complete Markdown from a small selected folder. Apply the parser below to one saved output at a time. For invoice tables, continue with the invoice data extraction workflow to define fields and reconcile totals before export.

Keep rows, headers and spans together

Simple rectangular tables become escaped GFM pipe tables. Complex tables with merged cells or nested content remain sanitized HTML inside Markdown. The dated invoice capture demonstrates both representations.

Check table headers, units, totals and spanning labels against the PDF. A complete response does not guarantee every extracted cell is correct. For local alternatives, pdfplumber supports text and explicit-edge strategies as well as gridlines; see the current tool guide.

Parsing simple GFM output into pandas

First save validated conversion output as document.md. Install pandas in your Python environment, then save this script as table_to_csv.py and run it with Python 3. It reads a local file and exports only the first valid table candidate as CSV to stdout. If that first candidate is complex HTML, or no table exists, it fails explicitly. Later tables are not exported.

from pathlib import Path
import pandas as pd

# BEGIN SYNCED GFM TABLE PARSER
from html import unescape

def split_gfm_row(line):
    row = line.strip()
    if not (row.startswith("|") and row.endswith("|")):
        raise ValueError("Expected a GFM table row with outer pipes")

    cells, cell = [], []
    escaped = False
    for char in row[1:-1]:
        if escaped:
            if char in ("\\", "|", "*", "_", "`"):
                cell.append(char)
            else:
                cell.extend(("\\", char))
            escaped = False
        elif char == "\\":
            escaped = True
        elif char == "|":
            cells.append(unescape("".join(cell).strip()))
            cell = []
        else:
            cell.append(char)
    if escaped:
        cell.append("\\")
    cells.append(unescape("".join(cell).strip()))
    return cells

def parse_gfm_candidate(lines):
    if len(lines) < 2:
        return None
    rows = [split_gfm_row(line) for line in lines]
    width = len(rows[0])
    if width == 0 or len(rows[1]) != width:
        return None
    if any(cell != "---" for cell in rows[1]):
        return None
    if any(len(row) != width for row in rows[2:]):
        return None
    return rows

def iter_table_candidates(markdown):
    current = []
    for line in markdown.splitlines():
        row = line.strip()
        if row.startswith("|") and row.endswith("|"):
            current.append(row)
            continue
        if current:
            candidate = parse_gfm_candidate(current)
            if candidate is not None:
                yield "gfm", candidate
            current = []
        if "<table" in row.casefold():
            yield "html", markdown
    if current:  # flush a table at end of output
        candidate = parse_gfm_candidate(current)
        if candidate is not None:
            yield "gfm", candidate

def extract_gfm_tables(markdown):
    return [value for kind, value in iter_table_candidates(markdown) if kind == "gfm"]

def first_table_branch(markdown):
    return next(iter_table_candidates(markdown), ("none", None))
# END SYNCED GFM TABLE PARSER

markdown = Path("document.md").read_text(encoding="utf-8")
table_kind, table_rows = first_table_branch(markdown)
if table_kind == "html":
    raise ValueError("Complex sanitized HTML table found: use a non-rendering HTML parser; do not render it as trusted HTML.")
if table_kind == "none":
    raise ValueError("No valid table found")
header = table_rows[0]
data = table_rows[2:]
df = pd.DataFrame(data, columns=header)
print(df.to_csv(index=False), end="")

The helper splits structural pipes before unescaping cells. It requires a header, canonical separator and consistent row widths. Treat cell text as data. For complex sanitized HTML, use a non-rendering HTML parser that retains span attributes; do not render API output as trusted HTML.

Convert a document

Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.

See the API reference for authentication, limits and recovery.

Optional API demo: invoice PDF to Markdown

Bash, curl and Python 3. This demo script validates success before printing Markdown.

#!/usr/bin/env bash
set -euo pipefail
key=demo_public_key
identity=demo-example
work=$(mktemp -d)
trap 'rm -rf "$work"' EXIT
result=$work/result
status=$(curl --silent --show-error --max-time 660 -o "$result" -w '%{http_code}' \
  "https://pdftomarkdown.dev/v1/convert" \
  -H "Authorization: Bearer $key" -H "Idempotency-Key: $identity" \
  -H "Content-Type: application/json" \
  --data '{"input":{"pdf_url":"https://pdftomarkdown.dev/samples/invoice.pdf"}}')
if [ "$status" != 200 ]; then echo "HTTP $status; keep the identity for recovery" >&2; exit 1; fi
python3 - "$result" <<'PY'
import json, sys
with open(sys.argv[1]) as response:
    result = json.load(response)
if not (isinstance(result, dict) and result.get("complete") is True
        and isinstance(result.get("markdown"), str)
        and type(result.get("pages")) is int and result["pages"] >= 0
        and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
    sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
PY
{
  "complete": true,
  "markdown": "# CONTOSO LTD.\n\n# INVOICE\n\nContoso Headquarters\n\n123 456th St\n\nNew York, NY, 10001\n\nINVOICE: INV-100\n\nDATE: 11/15/2019\n\nDUE DATE: 12/15/2019\n\nCUSTOMER NAME: MICROSOFT CORPORATION\n\nCUSTOMER ID: CID-12345\n\nMicrosoft Corp\n\n123 Other St,\n\nRedmond WA, 98052\n\nBILL TO:\n\nMicrosoft Finance\n\n123 Bill St,\n\nRedmond WA, 98052\n\nSHIP TO:\n\nMicrosoft Delivery\n\n123 Ship St,\n\nRedmond WA, 98052\n\nSERVICE ADDRESS:\n\nMicrosoft Services\n\n123 Service St,\n\nRedmond WA, 98052\n\n| SALESPERSON | P.O. NUMBER | REQUISITIONER | SHIPPED VIA | F.O.B. POINT | TERMS |\n| --- | --- | --- | --- | --- | --- |\n|  | PO-3333 |  |  |  |  |\n\n<table><tr><th>QUANTITY</th><th>DESCRIPTION</th><th>UNIT PRICE</th><th>TOTAL</th></tr><tr><td>1</td><td>Test for 23 fields</td><td>1</td><td>$100.00</td></tr><tr><td></td><td></td><td></td><td></td></tr><tr><td colspan=\"3\">SUBTOTAL</td><td>$100.00</td></tr><tr><td colspan=\"3\">SALES TAX</td><td>$10.00</td></tr><tr><td colspan=\"3\">TOTAL</td><td>$110.00</td></tr><tr><td colspan=\"3\">PREVIOUS BALANCE</td><td>$500.00</td></tr><tr><td colspan=\"3\">TOTAL DUE</td><td>$610.00</td></tr></table>\n\nTHANK YOU FOR YOUR BUSINESS!\n\nREMIT TO:\n\nContoso Billing\n\n123 Remit St\n\nNew York, NY, 10001\n\n> Processed by pdfToMarkdown.dev",
  "pages": 1,
  "request_id": "req_example_invoice"
}

Pricing

Try page 1 free, or get 20 trial pages once with a new account. Then choose paid credits. Monthly pages expire; top-ups do not.