Turn logistics PDFs into structured JSON
Structured fields from CMR, invoices, POD and more — digital PDFs and scans. One REST endpoint. 1 credit = 1 page. Failures refunded.
No credit card · Email verification · 60 req/min · API quickstart
curl -X POST https://cyber.lerepop.lv/api/v1/extract \
-H "Authorization: Bearer ext_live_YOUR_KEY" \
-F "file=@invoice.pdf" \
-F "document_type=invoice"
Example response
{
"status": "completed",
"document_type": "invoice",
"page_count": 1,
"credits_charged": 1,
"fields": {
"invoice_number": { "value": "INV-2026-0042", "confidence": 0.97 },
"amount": { "value": 1240.00, "confidence": 0.95 },
"currency": { "value": "EUR", "confidence": 0.99 }
}
}
—
Key fields correct
on our 90-doc test set
—
Invoice · CMR · POD
EU freight slice
0
Test documents
synthetic corpus
0
Document types
0
Languages
50 MB
Max upload
“Key fields correct” = how often we matched the shipment-critical fields (parties, refs, dates, amounts) on a fixed lab set — not a promise for every customer PDF. How we measure →
Interactive preview
See real extraction output
Pick a freight document — including multi-page CMR and combined PDFs. Click a field to highlight it on the page (and zoom to that spot). Hover does not change the selection.
Fields appear here after load.
Synthetic benchmark documents — no customer data. Live API: 1 credit = 1 page; failures refund. How we measure accuracy → · Sign up — try your PDF
Why Extract API
Production-grade extraction, developer-first billing
Sync or async. Digital PDFs parsed locally, scans via OCR. One credit per page — cached repeats are free.
Sync & async
Blocking responses for simple flows, or queue + webhook for heavy batches. Poll status anytime.
Hybrid pipeline
Poppler locally for digital PDFs. Mistral OCR for scans. Optional structured LLM — disclosed in our sub-processor list.
Credit-based billing
One credit = one PDF page. Cached repeats (same account) are free. Failed extractions are refunded. Buy packs via Stripe Checkout.
EU freight languages
English, German, Estonian, Lithuanian, Latvian, Russian, Polish — tuned for Baltic and Central European documents.
Transparent & GDPR-ready
Per-account cache isolation. DPA, privacy policy, sub-processors list. Full data-flow docs for your compliance team.
8 document types
Invoice, CMR, POD, delivery note, customs, bill of lading, rail waybill — or auto-detect with combined PDFs.
Document types
Built for freight paperwork
Pass document_type
or let the pipeline auto-detect. Every field comes back with a confidence score.
invoice
Invoice
Number, dates, amounts, VAT, parties
cmr
CMR
Consignor, consignee, goods, weight
pod
POD
Delivery date, signature, reference
delivery_note
Delivery note
Shipper, recipient, line items
customs
Customs
Declaration type, HS codes, values
bill_of_lading
Bill of lading
Vessel, ports, containers
rail_waybill
Rail waybill
Wagon, stations, cargo
auto
Auto / combined
Multi-page PDF split + classify
One endpoint
Same API call for every document type
-F "document_type=cmr"
Getting started
From signup to first request
Four short steps. No credit card — 15 free page credits after email verification.
Sign up on the web with email, password, and company name. Then open the verification link we send — free page credits unlock after verification.
Sign up free
After verification, the dashboard shows your full API key once
(prefix ext_live_…).
Store it in a secret manager — it cannot be retrieved again in full.
Auth header: Authorization: Bearer …
(not X-Api-Key).
POST /api/v1/extract
with a PDF and document_type.
One credit per page; failures are refunded.
curl -X POST https://cyber.lerepop.lv/api/v1/extract \
-H "Authorization: Bearer ext_live_YOUR_KEY" \
-F "file=@invoice.pdf" \
-F "document_type=invoice"
Prefer the OpenAPI explorer? API reference →
Each field returns value,
confidence, and often
source.
Use confidence thresholds before writing into your TMS.
- Check balance:
GET /api/v1/account/balance - Buy page-credit packs via Stripe when free quota is used
- Enable async + webhooks for batch volume
Honest fit
Where we are strong — and where to review
We optimise for freight cores (parties, refs, amounts, dates, MRN/BL IDs), not every JSON key.
Confidence scores and needs_review exist for a reason.
Strong today
Digital invoice & POD
Clean digital PDFs and light scans score well on our lab set. Best starting point for TMS invoice/POD automation.
Good with review
CMR & multi-page
Standard CMR layouts work well; ambiguous parties and hard scans still need confidence thresholds or human review on critical fields.
Pilot first
Customs, B/L, ugly scans
Supported in the API, but expect more review on stamps, occlusion, and exotic templates. Run a paid pilot on your PDFs before volume commit.
Numbers on this site are key-field match rates on a fixed lab set — not a promise for every customer PDF. See how we measure →
Trust & compliance
Clear about data — before you integrate
Freight PDFs contain personal data. We document exactly what we store, cache, and send to third parties.
Local first
Digital PDFs are parsed on our EEA servers with Poppler. No cloud AI unless the document needs OCR or structured extraction.
Sub-processors disclosed
Scans may use Mistral AI; optional stages may use OpenAI. Full list, triggers, and 30-day change notice — /subprocessors.
Per-account cache
Results cached 30 days in Redis — field values only, not PDFs. Never shared between API accounts. Details →
FAQ
Questions teams ask before integrating
Billing, accuracy, privacy, and limits — answered plainly.
One credit equals one PDF page. A 3-page CMR reserves 3 credits when the request is accepted. Cache hits and failed extractions do not keep the charge — failures are refunded.
No. After email verification you get 15 free page credits per month. Buy Stripe credit packs only when you need more volume.
We publish how often shipment-critical fields match a labelled lab set (shown as a %). That is not a guarantee for every customer scan. Digital invoices and PODs are strongest; hard scans and exotic customs/B/L need review more often. Run a short pilot on your own files.
Invoice, CMR, POD, delivery note, customs, bill of lading, rail waybill, plus auto-detect and combined multi-document PDFs. Pass document_type or let the pipeline classify.
We do not keep uploaded PDFs as a long-term archive for extraction cache. Field results may be cached per account for 30 days in Redis and are never shared across API accounts. See data-processing docs and the DPA for the full picture.
Digital PDFs are parsed locally (Poppler) on our servers. Scans may use Mistral OCR; optional structured stages may use OpenAI. Triggers and the full list live on /subprocessors with change notice.
Use sync POST /extract for interactive or low-volume flows. For batches, use async mode, poll request status, and optionally webhooks. Rate limit is 60 requests/minute by default.
Uploads up to 50 MB. Billable pages are capped at 20 per request — larger files return a clear validation error.
No. The interactive demo replays curated fixture PDFs and stored extraction results (with page highlights). Sign up and use the dashboard or API to run your own documents.
Still unsure? Read the docs or start free.
Pricing
Simple pricing, no surprises
1 credit = 1 PDF page. Cache hits and failed extractions do not keep the charge. Multi-page files reserve credits up front.