01
Two releases, two purposes
DocTorture v0.1 is the retained historical 1,000-document synthetic invoice OCR evaluation set. Its four-engine execution contains 4,000 verified sample evidence records. Public comparative views show Google native and Amazon OCR.
DocTorture-10K v1.0 is a separate sealed dataset release: 10,000 unique samples grouped by 2,500 sources, with four severity cohorts. Dataset and benchmark versions identify different artifacts. Published v0.1 model scores do not measure the v1.0 corpus, and sealing v1.0 does not replace the frozen v0.1 benchmark.
02
Source-group separation
Clean and degraded variants belonging to the same source stay in the same split. The 10K release assigns 8,000 samples to train, 1,000 to validation and 1,000 to test. Each severity cohort contains 2,500 samples.
This grouping prevents direct leakage of one source document across splits. It does not prove independence from all template or renderer characteristics; five source templates still define a deliberately bounded document distribution.
03
Admitted v1.0 scope
The eight idealized families are blur, motion blur, JPEG compression, perspective, rotation/skew, page cropping, contrast loss and color cast. Generation includes documented ordered mixtures.
Physical coffee, water and glare effects are excluded from calibrated commercial v1.0 scope.
04
Evaluation contract
Freeze the input manifest, truth sidecar, provider configuration and scoring policy before a run. Accept a sample only after its raw response, normalized result and evidence checksums are durably verified.
A complete comparison requires identical unique sample-ID sets across engines.