Privacy Policy
How pdfToMarkdown handles account data, usage logs, analytics, and PDF content.
Last updated: September 9, 2026
Service overview
pdfToMarkdown is a server-side PDF to Markdown API. Requests pass through Cloudflare for authentication, PDF inspection, and private temporary storage before the configured OCR processor handles the selected pages. Mistral OCR 4.1 is the current OCR processor.
Document content
PDF files and OCR output are processed to return the API response. Inspected source bytes are placed in private temporary object storage and are inaccessible through each signed OCR URL after 15 minutes. We request immediate deletion after a definite provider outcome; interrupted or ambiguous work is removed by bounded cleanup and storage lifecycle rules. A successful account response is temporarily cached as described below.
For signed-in conversions, an immutable private draft remains usable for 30 minutes, including while completing a purchase when Checkout is available. Starting conversion creates a separate temporary OCR copy with a signed URL valid for up to 15 minutes. Expired draft sources are removed by bounded cleanup; physical deletion can lag expiry. Successful account conversion responses, including Markdown, can be replayed for 24 hours without charging again. Access then expires; stored copies are removed by bounded cleanup and storage lifecycle rules, so physical deletion can occur later. Billing receipts and credit movements are retained separately from document content for accounting and payment reconciliation.
Account and API data
If you sign in with GitHub, we store the GitHub account ID, login, optional name, optional email, avatar URL, API key metadata, quota, and timestamps needed to issue and manage your developer key.
Usage logs
We log API usage metadata such as API key, tier, processed page count, status code, source IP, and timestamp. These logs help enforce quotas, diagnose reliability issues, and understand product usage. We do not need document text in analytics events.
Analytics
We use PostHog to measure website and API usage, including page views, CTA clicks, OAuth flow events, API request status, latency, and error codes. Analytics events must not include PDF content or converted Markdown.
Third-party processors
The service uses Cloudflare for edge hosting, routing, private temporary storage, and database infrastructure; Mistral for OCR 4.1 processing; GitHub for OAuth; and PostHog for analytics. Mistral receives a short-lived signed URL to our inspected copy, not the customer's original mutable URL. If you purchase credits or a subscription when sales are available, Stripe handles hosted Checkout and payment-related account data.
Your choices
You can avoid GitHub login by using the public demo key, subject to its limits. You can request account or usage-data deletion by contacting privacy@pdftomarkdown.dev.
Contact
For privacy questions, contact privacy@pdftomarkdown.dev.