Data processing & privacy
Extract API processes freight PDFs on your behalf. This page explains what data goes where, how caching works, which third-party providers may receive document content, and your GDPR obligations as a customer.
Roles under GDPR
| Data | You (customer) | Extract API |
|---|---|---|
| Your account email, company name, billing | — | Data controller |
| Personal data inside uploaded PDFs (names, addresses, amounts in CMR/invoices) | Data controller | Data processor (see DPA) |
By registering for an API key you accept our Data Processing Agreement. Full legal text: Privacy Policy, Sub-processors.
Processing pipeline — what leaves your infrastructure
PDF upload
│
├─ Digital PDF (embedded text layer)
│ └─ Poppler (pdftotext) on our servers — text stays in EEA hosting
│
├─ Scanned / image PDF (little or no text)
│ └─ Mistral OCR (if MISTRAL_OCR_ENABLED) — PDF sent to Mistral over TLS
│
├─ Structured field extraction (complex / ambiguous fields)
│ └─ Mistral or OpenAI JSON schema (if enabled) — extracted text or crops sent over TLS
│
└─ Vision follow-up (optional, rare fields)
└─ OpenAI Vision (if DOCUMENT_EXTRACTION_VISION_ENABLED) — page images over TLS
What stays on our servers only
- Digital PDFs with a readable text layer — processed locally with Poppler; no cloud OCR required.
- Rule-based parsers (regex, CMR box labels) — no third-party LLM.
- API access logs — method, path, status, duration; not document body content.
When third parties receive document content
| Provider | Trigger | Data sent | Region |
|---|---|---|---|
| Mistral AI | Scanned PDF, insufficient local text, or structured/VLM passes | PDF bytes or extracted text / image crops | EU preferred; confirm in your Mistral contract |
| OpenAI | Structured provider = openai, or vision enabled | Extracted text or page images | US — Standard Contractual Clauses |
| Stripe | Credit pack checkout | Email, company name, payment metadata — not PDFs | EU/US per Stripe |
Current sub-processor list and change-notice policy: /subprocessors. We notify account holders at least 30 days before adding or replacing a sub-processor.
Environment flags (operator / deployment)
Your Extract API operator controls which cloud stages are active:
| Variable | Default | Effect |
|---|---|---|
MISTRAL_OCR_ENABLED | false | Cloud OCR for scans |
MISTRAL_STRUCTURED_EXTRACTION_ENABLED | false | Mistral JSON extraction |
STRUCTURED_EXTRACTION_PROVIDER | mistral | mistral or openai |
DOCUMENT_EXTRACTION_VISION_ENABLED | false | OpenAI vision follow-up |
EXTRACT_CACHE_ENABLED | true | Result cache in Redis |
Enterprise deployments can run with cloud stages disabled (Poppler + rules only) for documents that have a text layer. Scanned documents would return degraded or empty fields without OCR enabled. Contact sales@extract.example.com for EU-only or on-prem options.
Result caching
To avoid re-processing identical documents, we cache extraction results (structured fields and confidence scores) in Redis. The original PDF file is not stored in the cache.
Cache key scope
- SHA-256 hash of PDF content
document_typeyou passed in the request- Your API account ID — cache entries are never shared between customers
What is stored in cache
- Extracted field values (may include personal data from the document)
- Confidence scores, detected document type, internal pipeline metadata
- Timestamp (
cached_at)
Cache behaviour for API clients
| First upload | Repeat (same account, same PDF, same type) | |
|---|---|---|
| Status | completed | cached |
| Credit charged | Yes | No |
| Processing time | Normal | ~0 ms |
| Cloud OCR / LLM called | If needed | No — served from cache |
Retention & deletion
- TTL: 30 days (configurable via
EXTRACT_CACHE_TTL_DAYS) - Uploaded PDFs: deleted after extraction (async: after the job completes)
- Extraction request records: stored in database for billing and support
- To request early deletion of cached results or account data: privacy@extract.example.com
Your responsibilities as a customer
- Ensure you have a lawful basis to upload documents and any personal data they contain.
- Inform your own data subjects where required (e.g. employees, carriers, consignees named in PDFs).
- Review our sub-processors before processing sensitive data.
- Use async webhooks over HTTPS only (see Webhooks).
- Do not upload unlawful content or documents you lack rights to process.
International transfers
Primary hosting is in the EEA. When Mistral or OpenAI process document content outside the EEA, we rely on Standard Contractual Clauses and vendor DPAs as described in our DPA Section 7.
Security contacts
- Privacy / data subject requests: privacy@extract.example.com
- DPA / sub-processor objections: dpa@extract.example.com
- Security incidents: security@extract.example.com