Ex Extract API

Quality

How good is extraction — in plain terms

We check whether the important shipment fields (parties, references, dates, amounts, route) match a labelled answer key. The headline % is that match rate on a fixed lab set — not a guarantee for every customer PDF.

Last updated 2026-08-19

—

Key fields correct

— documents · full mix

—

Invoice · CMR · POD

EU freight types only

—

Documents in live test

same set every run

What “key fields correct” means

Imagine a CMR must fill: number, consignee, carrier, origin, destination, pickup date. We compare each extracted value to the label. The score balances finding the right values and not inventing wrong ones (engineers call this F1). We report it as a percentage so buyers can read it at a glance.

Automated regression check

On every code push we re-run 0 labelled fixtures so quality does not silently regress. The digital replay score (100%) uses stored text — it is a safety gate, not the live GPT/OCR result above. Maturity work also uses a 0-document synthetic corpus (0 layouts).

Test group Documents Key-field score Missed key fields
digital 0 100% 0
all required 0 100% 0

Live lab run (engineers)

Full pipeline with live OCR/LLM on the same —-doc mix, cache off. Key fields = product score. Extra columns keep older all-field scoring for continuity — ignore them unless you are comparing historical gates.

Live audit JSON is not available on this deploy.

What we do not claim

  • Blanket “95%+ on any CMR / any phone photo” — hard unseen scans can drop sharply vs the lab set.
  • The automated CI replay % as production quality — that gate does not call live GPT/OCR.
  • Equal strength on every type — invoices/PODs (digital) are strongest; customs / B/L / ugly scans need more review.

Best next step: a short pilot on your own PDFs with an agreed field checklist. Start free →

Shipment-critical fields we score

These are what “key fields” means in the percentages above.

reference_number customer_name carrier_name carrier_cost carrier_cost_currency origin destination pickup_date

How the benchmark runs

  1. 1 Fixture replay regression gate — OCR text from `.txt` fixtures, structured LLM responses HTTP-faked. Runs in CI on every push on 0 labelled benchmark fixtures. Digital core weighted F1 of 1.0000 is this replay, not live extraction.
  2. 2 Live synthetic audit — GPT structured + Mistral OCR, cache off, 90-doc mix (60 pipeline + 30 maturity). Product metric is core F1 (CorpusCoreFieldMap). Legacy overall F1 (all-field pipeline) is shown for continuity. Both include zeros. Neither is 1.0000.
  3. 3 0-document HTML-rendered synthetic corpus (0 types, 0 languages, 0 layouts) for maturity testing — labels only, no customer PDFs.
  4. 4 Scan degradation axes (visual noise, OCR noise, completeness) on synthetic scan PDFs — tracked in scan_hard and invoice_scan benchmark groups.
  5. 5 Core fields are the shipment-critical fields used in freight automation (reference, parties, route, dates, carrier cost).

CI gate

Every push runs the digital fixture benchmark plus API feature tests (18 tests). Minimum key-field score: 95%.

No live keys in CI

Fixture replay uses stored OCR text and faked structured responses — reproducible, fast, no Mistral/OpenAI spend in that job.

0 types · 0 languages · 0 layouts

Invoice, CMR, POD, delivery note, customs, B/L, rail waybill, combined PDFs — English, German, Estonian, Lithuanian, Latvian, Russian, Polish. Synthetic corpus with scan degradation axes.