Ex Extract API
Freight document extraction API

Turn logistics PDFs into structured JSON

Structured fields from CMR, invoices, POD and more — digital PDFs and scans. One REST endpoint. 1 credit = 1 page. Failures refunded.

No credit card · Email verification · 60 req/min · API quickstart

curl -X POST https://cyber.lerepop.lv/api/v1/extract \
  -H "Authorization: Bearer ext_live_YOUR_KEY" \
  -F "file=@invoice.pdf" \
  -F "document_type=invoice"

Example response

{
  "status": "completed",
  "document_type": "invoice",
  "page_count": 1,
  "credits_charged": 1,
  "fields": {
    "invoice_number": { "value": "INV-2026-0042", "confidence": 0.97 },
    "amount":         { "value": 1240.00, "confidence": 0.95 },
    "currency":       { "value": "EUR", "confidence": 0.99 }
  }
}

—

Key fields correct

on our 90-doc test set

—

Invoice · CMR · POD

EU freight slice

0

Test documents

synthetic corpus

0

Document types

0

Languages

50 MB

Max upload

“Key fields correct” = how often we matched the shipment-critical fields (parties, refs, dates, amounts) on a fixed lab set — not a promise for every customer PDF. How we measure →

Interactive preview

See real extraction output

Pick a freight document — including multi-page CMR and combined PDFs. Click a field to highlight it on the page (and zoom to that spot). Hover does not change the selection.

Loading preview…

Source document
100%
—
Loading document preview…
API fields

Fields appear here after load.

Synthetic benchmark documents — no customer data. Live API: 1 credit = 1 page; failures refund. How we measure accuracy → · Sign up — try your PDF

Why Extract API

Production-grade extraction, developer-first billing

Sync or async. Digital PDFs parsed locally, scans via OCR. One credit per page — cached repeats are free.

Sync & async

Blocking responses for simple flows, or queue + webhook for heavy batches. Poll status anytime.

Hybrid pipeline

Poppler locally for digital PDFs. Mistral OCR for scans. Optional structured LLM — disclosed in our sub-processor list.

Credit-based billing

One credit = one PDF page. Cached repeats (same account) are free. Failed extractions are refunded. Buy packs via Stripe Checkout.

EU freight languages

English, German, Estonian, Lithuanian, Latvian, Russian, Polish — tuned for Baltic and Central European documents.

Transparent & GDPR-ready

Per-account cache isolation. DPA, privacy policy, sub-processors list. Full data-flow docs for your compliance team.

8 document types

Invoice, CMR, POD, delivery note, customs, bill of lading, rail waybill — or auto-detect with combined PDFs.

Document types

Built for freight paperwork

Pass document_type or let the pipeline auto-detect. Every field comes back with a confidence score.

Full field reference →
invoice

Invoice

Number, dates, amounts, VAT, parties

cmr

CMR

Consignor, consignee, goods, weight

pod

POD

Delivery date, signature, reference

delivery_note

Delivery note

Shipper, recipient, line items

customs

Customs

Declaration type, HS codes, values

bill_of_lading

Bill of lading

Vessel, ports, containers

rail_waybill

Rail waybill

Wagon, stations, cargo

auto

Auto / combined

Multi-page PDF split + classify

One endpoint

Same API call for every document type

-F "document_type=cmr"

Getting started

From signup to first request

Four short steps. No credit card — 15 free page credits after email verification.

Sign up on the web with email, password, and company name. Then open the verification link we send — free page credits unlock after verification.

Sign up free

Honest fit

Where we are strong — and where to review

We optimise for freight cores (parties, refs, amounts, dates, MRN/BL IDs), not every JSON key. Confidence scores and needs_review exist for a reason.

Strong today

Digital invoice & POD

Clean digital PDFs and light scans score well on our lab set. Best starting point for TMS invoice/POD automation.

Good with review

CMR & multi-page

Standard CMR layouts work well; ambiguous parties and hard scans still need confidence thresholds or human review on critical fields.

Pilot first

Customs, B/L, ugly scans

Supported in the API, but expect more review on stamps, occlusion, and exotic templates. Run a paid pilot on your PDFs before volume commit.

Numbers on this site are key-field match rates on a fixed lab set — not a promise for every customer PDF. See how we measure →

Trust & compliance

Clear about data — before you integrate

Freight PDFs contain personal data. We document exactly what we store, cache, and send to third parties.

Local first

Digital PDFs are parsed on our EEA servers with Poppler. No cloud AI unless the document needs OCR or structured extraction.

Sub-processors disclosed

Scans may use Mistral AI; optional stages may use OpenAI. Full list, triggers, and 30-day change notice — /subprocessors.

Per-account cache

Results cached 30 days in Redis — field values only, not PDFs. Never shared between API accounts. Details →

FAQ

Questions teams ask before integrating

Billing, accuracy, privacy, and limits — answered plainly.

One credit equals one PDF page. A 3-page CMR reserves 3 credits when the request is accepted. Cache hits and failed extractions do not keep the charge — failures are refunded.

Still unsure? Read the docs or start free.

Pricing

Simple pricing, no surprises

1 credit = 1 PDF page. Cache hits and failed extractions do not keep the charge. Multi-page files reserve credits up front.