Quality
How good is extraction — in plain terms
We check whether the important shipment fields (parties, references, dates, amounts, route) match a labelled answer key. The headline % is that match rate on a fixed lab set — not a guarantee for every customer PDF.
Last updated 2026-08-19
—
Key fields correct
— documents · full mix
—
Invoice · CMR · POD
EU freight types only
—
Documents in live test
same set every run
What “key fields correct” means
Imagine a CMR must fill: number, consignee, carrier, origin, destination, pickup date. We compare each extracted value to the label. The score balances finding the right values and not inventing wrong ones (engineers call this F1). We report it as a percentage so buyers can read it at a glance.
Automated regression check
On every code push we re-run 0 labelled fixtures so quality does not silently regress. The digital replay score (100%) uses stored text — it is a safety gate, not the live GPT/OCR result above. Maturity work also uses a 0-document synthetic corpus (0 layouts).
| Test group | Documents | Key-field score | Missed key fields |
|---|---|---|---|
| digital | 0 | 100% | 0 |
| all required | 0 | 100% | 0 |
Live lab run (engineers)
Full pipeline with live OCR/LLM on the same —-doc mix, cache off. Key fields = product score. Extra columns keep older all-field scoring for continuity — ignore them unless you are comparing historical gates.
Live audit JSON is not available on this deploy.
What we do not claim
- Blanket “95%+ on any CMR / any phone photo” — hard unseen scans can drop sharply vs the lab set.
- The automated CI replay % as production quality — that gate does not call live GPT/OCR.
- Equal strength on every type — invoices/PODs (digital) are strongest; customs / B/L / ugly scans need more review.
Best next step: a short pilot on your own PDFs with an agreed field checklist. Start free →
Shipment-critical fields we score
These are what “key fields” means in the percentages above.
reference_number
customer_name
carrier_name
carrier_cost
carrier_cost_currency
origin
destination
pickup_date
How the benchmark runs
- 1 Fixture replay regression gate — OCR text from `.txt` fixtures, structured LLM responses HTTP-faked. Runs in CI on every push on 0 labelled benchmark fixtures. Digital core weighted F1 of 1.0000 is this replay, not live extraction.
- 2 Live synthetic audit — GPT structured + Mistral OCR, cache off, 90-doc mix (60 pipeline + 30 maturity). Product metric is core F1 (CorpusCoreFieldMap). Legacy overall F1 (all-field pipeline) is shown for continuity. Both include zeros. Neither is 1.0000.
- 3 0-document HTML-rendered synthetic corpus (0 types, 0 languages, 0 layouts) for maturity testing — labels only, no customer PDFs.
- 4 Scan degradation axes (visual noise, OCR noise, completeness) on synthetic scan PDFs — tracked in scan_hard and invoice_scan benchmark groups.
- 5 Core fields are the shipment-critical fields used in freight automation (reference, parties, route, dates, carrier cost).
CI gate
Every push runs the digital fixture benchmark plus API feature tests (18 tests). Minimum key-field score: 95%.
No live keys in CI
Fixture replay uses stored OCR text and faked structured responses — reproducible, fast, no Mistral/OpenAI spend in that job.
0 types · 0 languages · 0 layouts
Invoice, CMR, POD, delivery note, customs, B/L, rail waybill, combined PDFs — English, German, Estonian, Lithuanian, Latvian, Russian, Polish. Synthetic corpus with scan degradation axes.