Turn benchmark results into claims that hold up.

Your team knows what went into the result. Everyone else has to take your word for it. Colophon closes that gap with a public claim people can inspect for themselves.

A real published claimSelf-run

Judging the LoCoMo judges

A controlled benchmark of the grading prompts behind published LoCoMo scores. Six grading configurations judged the same 240 answers on the same model snapshot.

What it found

Changing only the grader moved agreement with the same labels from 60.8% to 87.9%.

4,320 of 4,320planned calls accounted for

For results that need to leave the room.

A benchmark becomes valuable when people beyond its authors can rely on it.

You are publishing a benchmark

Give the result a permanent place people can cite, inspect, and return to.

You are making a performance claim

Give customers, reviewers, and competitors something stronger than your word.

You are relying on a result

Put a clear public record behind a product, research, or procurement decision.

Choose who runs the benchmark.

Run it yourself, or have Colophon run it. The published claim says which.

Self-run

Strong for most claims

Keep the run with your team while giving readers a result they can inspect for themselves.

Colophon-run

When added distance matters

Colophon runs your locked benchmark and seals the result, so the claim names who ran it.

Anyone can check the claim.

Checking is free. It stays free.

Get the files from the LoCoMo report
npx @colophon-claims/verify@0.2.1 ./bundle

The checker is available on npm.

Where does your claim need to go?

Tell us what you are trying to establish and who needs to rely on it. We can talk through which path fits.

Start a conversation