Extract documents
Upload a PDF and receive structured fields with confidence scores.
Endpoint
POST https://cyber.lerepop.lv/api/v1/extract
Content-Type: multipart/form-data
Authorization: Bearer ext_live_...
Form fields
| Field | Type | Required | Description |
|---|---|---|---|
file |
PDF binary | Yes | Max 50 MB. Must be application/pdf. |
document_type |
string | No | Default: auto. See document types. |
webhook_url |
URL | No | HTTPS URL for async completion callback. Required when options[async]=true. |
options[confidence_threshold] |
float 0–1 | No | Default: 0.7. Fields below threshold may be omitted or flagged. |
options[include_raw_text] |
boolean | No | Include extracted plain text in response (may increase payload size). |
options[async] |
boolean | No | Queue extraction and return 202. See async processing. |
Sync response (200)
{
"request_id": "uuid",
"status": "completed",
"document_type": "invoice",
"cached": false,
"processing_time_ms": 1840,
"credit_charged": true,
"fields": {
"field_key": {
"value": "extracted value",
"confidence": 0.95,
"source": "rule|structured|vision|ocr"
}
},
"warnings": [],
"business_rule_violations": []
}
Field structure
Each field in fields contains:
value— string, number, or nested object for composite fieldsconfidence— 0.0 to 1.0source— which pipeline layer produced the value
Caching
Identical PDF content (SHA-256 hash) + same document_type + same API account
returns a cached result:
cached: true,status: cachedprocessing_time_ms: 0credit_charged: false— no credits consumed (failure or cache)credits_charged— number of page credits billed (equalspage_counton success)page_count— PDF pages used for billing- No repeat calls to cloud OCR/LLM providers
Cache TTL: 30 days. Cached data is per-account only — never shared between customers. See Data processing — caching for GDPR details.
Credits
One credit is charged per PDF page for successful, non-cached extractions. Failures refund the reserved credits; cache hits do not charge.
Insufficient balance returns 402 Payment Required.
Example: CMR extraction
curl -X POST https://cyber.lerepop.lv/api/v1/extract \
-H "Authorization: Bearer ext_live_..." \
-F "file=@cmr-scan.pdf" \
-F "document_type=cmr" \
-F "options[confidence_threshold]=0.75"
Example: auto-detect combined PDF
curl -X POST https://cyber.lerepop.lv/api/v1/extract \
-H "Authorization: Bearer ext_live_..." \
-F "file=@combined-shipment.pdf" \
-F "document_type=combined"
Multi-page PDFs may be split and classified per page before field extraction.