Ex Extract API

Data processing & privacy

Extract API processes freight PDFs on your behalf. This page explains what data goes where, how caching works, which third-party providers may receive document content, and your GDPR obligations as a customer.

Roles under GDPR

DataYou (customer)Extract API
Your account email, company name, billing — Data controller
Personal data inside uploaded PDFs (names, addresses, amounts in CMR/invoices) Data controller Data processor (see DPA)

By registering for an API key you accept our Data Processing Agreement. Full legal text: Privacy Policy, Sub-processors.

Processing pipeline — what leaves your infrastructure

PDF upload
  │
  ├─ Digital PDF (embedded text layer)
  │    └─ Poppler (pdftotext) on our servers — text stays in EEA hosting
  │
  ├─ Scanned / image PDF (little or no text)
  │    └─ Mistral OCR (if MISTRAL_OCR_ENABLED) — PDF sent to Mistral over TLS
  │
  ├─ Structured field extraction (complex / ambiguous fields)
  │    └─ Mistral or OpenAI JSON schema (if enabled) — extracted text or crops sent over TLS
  │
  └─ Vision follow-up (optional, rare fields)
       └─ OpenAI Vision (if DOCUMENT_EXTRACTION_VISION_ENABLED) — page images over TLS

What stays on our servers only

  • Digital PDFs with a readable text layer — processed locally with Poppler; no cloud OCR required.
  • Rule-based parsers (regex, CMR box labels) — no third-party LLM.
  • API access logs — method, path, status, duration; not document body content.

When third parties receive document content

ProviderTriggerData sentRegion
Mistral AI Scanned PDF, insufficient local text, or structured/VLM passes PDF bytes or extracted text / image crops EU preferred; confirm in your Mistral contract
OpenAI Structured provider = openai, or vision enabled Extracted text or page images US — Standard Contractual Clauses
Stripe Credit pack checkout Email, company name, payment metadata — not PDFs EU/US per Stripe

Current sub-processor list and change-notice policy: /subprocessors. We notify account holders at least 30 days before adding or replacing a sub-processor.

Environment flags (operator / deployment)

Your Extract API operator controls which cloud stages are active:

VariableDefaultEffect
MISTRAL_OCR_ENABLEDfalseCloud OCR for scans
MISTRAL_STRUCTURED_EXTRACTION_ENABLEDfalseMistral JSON extraction
STRUCTURED_EXTRACTION_PROVIDERmistralmistral or openai
DOCUMENT_EXTRACTION_VISION_ENABLEDfalseOpenAI vision follow-up
EXTRACT_CACHE_ENABLEDtrueResult cache in Redis

Enterprise deployments can run with cloud stages disabled (Poppler + rules only) for documents that have a text layer. Scanned documents would return degraded or empty fields without OCR enabled. Contact sales@extract.example.com for EU-only or on-prem options.

Result caching

To avoid re-processing identical documents, we cache extraction results (structured fields and confidence scores) in Redis. The original PDF file is not stored in the cache.

Cache key scope

  • SHA-256 hash of PDF content
  • document_type you passed in the request
  • Your API account ID — cache entries are never shared between customers

What is stored in cache

  • Extracted field values (may include personal data from the document)
  • Confidence scores, detected document type, internal pipeline metadata
  • Timestamp (cached_at)

Cache behaviour for API clients

First uploadRepeat (same account, same PDF, same type)
Statuscompletedcached
Credit chargedYesNo
Processing timeNormal~0 ms
Cloud OCR / LLM calledIf neededNo — served from cache

Retention & deletion

  • TTL: 30 days (configurable via EXTRACT_CACHE_TTL_DAYS)
  • Uploaded PDFs: deleted after extraction (async: after the job completes)
  • Extraction request records: stored in database for billing and support
  • To request early deletion of cached results or account data: privacy@extract.example.com

Your responsibilities as a customer

  • Ensure you have a lawful basis to upload documents and any personal data they contain.
  • Inform your own data subjects where required (e.g. employees, carriers, consignees named in PDFs).
  • Review our sub-processors before processing sensitive data.
  • Use async webhooks over HTTPS only (see Webhooks).
  • Do not upload unlawful content or documents you lack rights to process.

International transfers

Primary hosting is in the EEA. When Mistral or OpenAI process document content outside the EEA, we rely on Standard Contractual Clauses and vendor DPAs as described in our DPA Section 7.

Security contacts

  • Privacy / data subject requests: privacy@extract.example.com
  • DPA / sub-processor objections: dpa@extract.example.com
  • Security incidents: security@extract.example.com