Extraction / by Understand Tech Lab
Turn documents into structured data
Compare an on-premises extraction workflow against a labeled set of documents and a fixed output schema.
Untested evaluation templateVersion 0.1.0No lab-verified results
Before running
Make the setup exact.
Model: Select text or vision model according to input documents; pin exact revision and precision.
Stack: OCR if required + local model server + schema validator. Pin each component.
Record your hardware SKU, memory allocation, topology, dataset revision and thresholds. The downloaded manifest is a starting specification; it does not install software.
Platform documentationEvaluation plan
- Prepare 30 shareable documents spanning formats, quality and layouts. Write a reference JSON output for each.
- Define required fields, types and treatment of missing values. Reserve some documents for a held-out test.
- Select an OS-compatible runtime and OCR pipeline if needed. Record model, prompt, preprocessing and dependency versions.
- Run the held-out documents without changing the prompt. Save raw outputs, validation errors, retries and human corrections.
- Repeat at the expected batch size and concurrency. Compare total completion time and energy per accepted document.
What counts as useful?
- Measure field-level accuracy and valid-schema rate separately.
- Count corrections and retries in cost and completion time.
- Set the acceptable error rate before reviewing held-out results.
Known limits
Document content and OCR quality can change results substantially.
Candidate machines are not validated deployments or capacity promises.