PDFをMarkdownに変換するAPI
エンドポイントは1つだけ。PDFを送ると、構造化されたMarkdownが返ってきます。 表・見出し・リストはそのまま保持。日本語のスキャン文書にも強い。登録不要で試せます。
Bash・curl・Python 3が必要です。成功レスポンスを検証してMarkdownを出力します。
#!/usr/bin/env bash
set -euo pipefail
key=demo_public_key
identity=demo-example
work=$(mktemp -d)
trap 'rm -rf "$work"' EXIT
result=$work/result
status=$(curl --silent --show-error --max-time 660 -o "$result" -w '%{http_code}' \
"https://pdftomarkdown.dev/v1/convert" \
-H "Authorization: Bearer $key" -H "Idempotency-Key: $identity" \
-H "Content-Type: application/json" \
--data '{"input":{"pdf_url":"https://pdftomarkdown.dev/samples/invoice.pdf"}}')
if [ "$status" != 200 ]; then echo "HTTP $status; keep the identity for recovery" >&2; exit 1; fi
python3 - "$result" <<'PY'
import json, sys
with open(sys.argv[1]) as response:
result = json.load(response)
if not (isinstance(result, dict) and result.get("complete") is True
and isinstance(result.get("markdown"), str)
and type(result.get("pages")) is int and result["pages"] >= 0
and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
PY{
"complete": true,
"markdown": "# CONTOSO LTD.\n\n# INVOICE\n\nContoso Headquarters\n\n123 456th St\n\nNew York, NY, 10001\n\nINVOICE: INV-100\n\nDATE: 11/15/2019\n\nDUE DATE: 12/15/2019\n\nCUSTOMER NAME: MICROSOFT CORPORATION\n\nCUSTOMER ID: CID-12345\n\nMicrosoft Corp\n\n123 Other St,\n\nRedmond WA, 98052\n\nBILL TO:\n\nMicrosoft Finance\n\n123 Bill St,\n\nRedmond WA, 98052\n\nSHIP TO:\n\nMicrosoft Delivery\n\n123 Ship St,\n\nRedmond WA, 98052\n\nSERVICE ADDRESS:\n\nMicrosoft Services\n\n123 Service St,\n\nRedmond WA, 98052\n\n| SALESPERSON | P.O. NUMBER | REQUISITIONER | SHIPPED VIA | F.O.B. POINT | TERMS |\n| --- | --- | --- | --- | --- | --- |\n| | PO-3333 | | | | |\n\n<table><tr><th>QUANTITY</th><th>DESCRIPTION</th><th>UNIT PRICE</th><th>TOTAL</th></tr><tr><td>1</td><td>Test for 23 fields</td><td>1</td><td>$100.00</td></tr><tr><td></td><td></td><td></td><td></td></tr><tr><td colspan=\"3\">SUBTOTAL</td><td>$100.00</td></tr><tr><td colspan=\"3\">SALES TAX</td><td>$10.00</td></tr><tr><td colspan=\"3\">TOTAL</td><td>$110.00</td></tr><tr><td colspan=\"3\">PREVIOUS BALANCE</td><td>$500.00</td></tr><tr><td colspan=\"3\">TOTAL DUE</td><td>$610.00</td></tr></table>\n\nTHANK YOU FOR YOUR BUSINESS!\n\nREMIT TO:\n\nContoso Billing\n\n123 Remit St\n\nNew York, NY, 10001\n\n> Processed by pdfToMarkdown.dev",
"pages": 1,
"request_id": "req_example_invoice"
}日本語文書のために選ぶ理由
現在の処理モデルはMistral OCR 4.1です。日付付きの入力PDFと変換結果を確認できます。2026年7月のCJKベンチマークは旧PaddleOCR処理系の履歴です。 日本語の文字誤り率0.10%(Tesseract 5.5.1: 0.99%)は、その模擬スキャンの実測値であり、現在のモデルの精度保証ではありません。
スキャンPDF・画像PDF
テキストレイヤーのない紙のスキャンでも、日本語の本文・見出しを認識してMarkdown化します。
複雑な表
単純な長方形の表はエスケープ済みGFMパイプテーブルに変換し、行・列の結合や入れ子を含む複雑な表は構造を壊さないようサニタイズ済みHTMLとしてMarkdown内に保持します。
段組みレイアウト
論文や報告書の多段組みでも、正しい読み順でテキストを再構成します。
対応している文書
請求書・領収書
研究論文(数式対応)
契約書
決算・財務資料
技術マニュアル
スキャン文書
Pythonから使う
HTTPリクエスト1回だけ。requests以外のインストールは不要です。
python
import base64
import os
import sys
from pathlib import Path
import requests
key = "demo_public_key"
identity = "demo-example"
if not key or not identity:
sys.exit("Set PDFTOMARKDOWN_API_KEY and a saved PDFTOMARKDOWN_IDEMPOTENCY_KEY")
try:
response = requests.post(
"https://pdftomarkdown.dev/v1/convert",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": identity,
"Content-Type": "application/json"},
json={"input": {"pdf_url": "https://pdftomarkdown.dev/samples/invoice.pdf"}},
timeout=(10, 660),
)
if response.status_code != 200:
sys.exit(f"HTTP {response.status_code}; keep the identity for recovery")
result = response.json()
if not (isinstance(result, dict) and result.get("complete") is True
and isinstance(result.get("markdown"), str)
and type(result.get("pages")) is int and result["pages"] >= 0
and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
except (requests.RequestException, ValueError, OSError):
sys.exit("Conversion failed; keep the identity and original input for recovery")コードを書かずに使う
公式CLIなら、ローカルファイル・URL・標準入力を1コマンドで変換できます。Node 18+があればインストール不要です。
terminal
$ npx pdftomarkdown 書類.pdf > 書類.mdCodexやClaude CodeなどのコーディングエージェントにPDFを読ませる方法は、CodexでPDFを読む方法をご覧ください。 詳細はCLIガイド(英語)へ。
ターミナルに貼り付けるだけ。アカウントは不要です。
#!/usr/bin/env bash
set -euo pipefail
key=demo_public_key
identity=demo-example
work=$(mktemp -d)
trap 'rm -rf "$work"' EXIT
result=$work/result
status=$(curl --silent --show-error --max-time 660 -o "$result" -w '%{http_code}' \
"https://pdftomarkdown.dev/v1/convert" \
-H "Authorization: Bearer $key" -H "Idempotency-Key: $identity" \
-H "Content-Type: application/json" \
--data '{"input":{"pdf_url":"https://pdftomarkdown.dev/samples/invoice.pdf"}}')
if [ "$status" != 200 ]; then echo "HTTP $status; keep the identity for recovery" >&2; exit 1; fi
python3 - "$result" <<'PY'
import json, sys
with open(sys.argv[1]) as response:
result = json.load(response)
if not (isinstance(result, dict) and result.get("complete") is True
and isinstance(result.get("markdown"), str)
and type(result.get("pages")) is int and result["pages"] >= 0
and isinstance(result.get("request_id"), str) and result["request_id"].strip()):
sys.exit("Incomplete or invalid conversion response")
sys.stdout.write(result["markdown"])
PY