# Benchmark comparison worksheet

Blank working document. Replace placeholders with your own evidence; an empty field is unknown, not a passing result.

Owner: [fill in]
Source revision: [fill in]
Environment: [fill in]
Evidence date: [fill in]

## Population

Claim under review: ___
Target population: ___
Selection rule: ___
Snapshot URLs/dates/hashes/licences: ___

## Identity and missingness

| Exact model ID | Source A identity/score | Source B identity/score | Join evidence or missing reason |
| --- | --- | --- | --- |
| ___ | unknown | unknown | ___ |

## Computation

Input JSON path/hash: ___
Matched/total/missing counts: not computed
Checker output/exit: not run
Sensitivity cohort and result: ___

## Interpretation

What the observed cohort supports: ___
What it cannot support: ___
Interval-comparison evidence, if any: ___
Review decision: ___

## Boundaries

- A restricted or frontier-only cohort does not represent all models. Correlation does not demonstrate task interchangeability or causation.

- Overlapping individual intervals do not prove equivalence. The checker produces no significance test and does not verify dataset rights or score provenance.

Workflow: https://undominated.ai/workflows/#audit-a-benchmark-comparison

Original worksheet: MIT. Linked resources retain their own licences and setup requirements.
