Retain per-case input/output identity and execution evidence.
Benchmark the exact model, not just the model name.
Evaluation becomes trustworthy when the benchmark definition, artifact subject, case outputs and result identity are connected. Zippri uses a worker-owned benchmark lifecycle so results are executed, retained and cryptographically sealed instead of being manually typed into a status field.
Reproducible evaluation definitions
Benchmark suites use a constrained evaluator contract so the definition of what was tested can be retained and hashed. This reduces ambiguity between repeated runs.
Cryptographically bind metrics to the exact run subject.
Rank only results that remain current and cryptographically valid.
Worker-owned execution
Queued runs are claimed by the benchmark executor, bound to the current immutable subject and completed by the worker. Manual lifecycle shortcuts are blocked once the benchmark stage is authoritative.
Current signed results
Result evidence includes the suite, exact artifact subject, case outputs, score and executor identity. A changed subject can make an old result stale rather than allowing it to remain current by label.
Common questions about AI Benchmarking
Why bind a benchmark to an immutable revision?
Otherwise a benchmark score can outlive the model bytes or configuration it actually measured.
Can benchmark results be entered manually?
The accepted Benchmark Cloud lifecycle is worker-owned; manual completion is intentionally blocked for authoritative benchmark runs.
What is the difference between evaluation and verification?
Evaluation measures model behavior or quality under a benchmark. Verification is a governed trust decision about evidence and approved scopes.