PDF Structure: Why Text Extraction Is Hard
npx pdftomarkdown your.pdfconverts page 1 of any PDF, key-free.Docs →PDF appearance and document structure are different
A PDF can describe glyph positions and page graphics without providing the logical reading order your application needs. Some PDFs also contain tags, form fields and other structure; inspect the actual file instead of assuming all PDFs lack semantics.
Text and font mappings
A rendered glyph and its Unicode text are different representations. Missing or incorrect mappings can make extracted text diverge from what a reader sees. OCR may help with visible content, but its transcription also needs checking.
Reading order and tables
Columns, sidebars, headings and merged table cells may require layout interpretation. Tagged structure can help where present. A plain text dump is only one extraction mode, not the full capability of every PDF library.
Scanned pages
An image-only page has no text layer to extract directly. PyMuPDF provides integrated OCR through separately installed Tesseract; see its OCR recipe, checked 10 September 2026.
Forms and signatures
Form-field values and visible page content may need separate handling. OCR does not verify a digital signature or document authenticity. Preserve the original PDF when those properties matter.
Choose an extraction path
Try direct text or structured extraction where it fits. Use OCR for scans and evaluate layout recovery when tables or reading order matter. The current tool guide distinguishes available capabilities from optional setup.
Convert a document
Use the tested Python, JavaScript or curl workflow. Choose one input: a public PDF URL, Base64 bytes, or a raw PDF upload. Raw uploads avoid Base64 expansion; all three send the document to an external processor.
The public demo converts page 1 with a watermark, at 3 requests per minute per IP. New accounts receive 20 trial pages once. Further account conversions use paid credits: monthly pages expire, top-ups do not. The API reference owns authentication, limits, replay and retention.