Menu
research
The engineering notebook.
Technical notes on generation, evaluation and the limits of the evidence.
Each note states what has been measured, what has been generated and what remains an open question.
DocTorture methodology
Dataset scope, source grouping and the boundary between generated inputs and measured model behavior.
02Why clean OCR benchmarks are insufficient
A clean baseline measures a useful boundary. It does not describe the full capture distribution.
03Measuring robustness across degradation severity
Keep recognition, serialization and cohort composition visible.
04Synthetic data for document AI training
Use a measured failure analysis to propose a training corpus, then test the result.
05Failure-conditioned dataset generation
Turn a failure map into an explicit, reproducible generation specification.
06Dataset provenance and reproducibility
Keep inputs, labels, recipes and evaluation evidence connected.
Build with evidence
See what breaks
before production does.
Run StressBench against your document model and get a condition-level robustness readout.