Product / Model robustness

Find where recognition
stops holding up.

A controlled evaluation of your document model, with condition-level diagnostics and retained evidence behind every accepted result.

Stress-test. Diagnose. Generate. Improve. Retest.

A repeatable evaluation loop.

Start from a controlled baseline.

Run your model or API against a fixed evaluation set. Preserve the original response alongside its input and configuration.

Measure recognition.
Inspect serialization.

Order-independent token F1 is the primary Track A recognition metric. Numeric-token and date-text accuracy show additional retention views. CER and WER remain separate reading-order-sensitive metrics.

A high sequence error alone does not invalidate strong token recognition. Preserve the raw output and frozen truth order, then inspect the source of disagreement.

Read the metric policy
0255075100Token F1 (%)CleanSubtleModerateSevere
Google Document AIAmazon OCR

Measured on DocTorture-1K v0.1 (1,000 documents), separate from the 10K v1.0 dataset. Unpaired synthetic cohort means. Clean n=250, subtle n=182, moderate n=147, severe n=421.

MODEL ROBUSTNESS REPORTIllustrative product UI

From a weak spot
to a testable next step.

01 / BaselineClean + severity curve
02 / DiagnosisWeakest conditions
03 / InspectionRepresentative failures
04 / RecognitionNumeric / date retention

RECOMMENDED TRAINING MIX · EXAMPLE

Motion blur35%
Crop25%
Perspective20%
Compression20%

Your recommendation follows your model’s measured failures.

Leave with a data decision.

An evaluation report connects the clean baseline and severity curve to weak conditions and representative failed inputs. A proposed training mix makes the next data experiment explicit.

The report interface shown here is illustrative.

Explore enterprise engagements

Build with evidence

See what breaks
before production does.

Run StressBench against your document model and get a condition-level robustness readout.