Claim

Judging the LoCoMo judges

Lowest to highest of the six grading prompts

60.8% to 87.9%

146 of 240 to 211 of 240 answers, agreement with the same correctness labels.

Changing only the grader.

  • Design posted
  • Sealed
  • Timestamp proof pending, as the proof states
  • Self-run
  • 4,320 of 4,320 planned calls accounted for

On the board: LoCoMo judge report frozen bank

What exactly ran

Who chose the tasks
The claimant chose them. Not a fixed public set, not drawn by lottery. Questions sampled by the claimant's seeded script; a model screened a 664-item pool, the operator hand-reviewed 255, and 137 were excluded and replaced to keep the 80/80/80 balance.
Method
LoCoMo judge report frozen bank 1.0.0, written by the claimant. All 240 planned items, 4,320 planned calls.
Arms
Six grading prompts, supplied by the claimant. Colophon does not look inside them.
Scorer
The method's own, fixed in the sealed run: gpt-4o-mini-2024-07-18 at temperature 0.0, through inspect-ai-judge 1. The majority of 3 calls decides each answer.

Who ran it

Self-run

The claimant ran it. The sealed disclosure opens:

“This is a local, self-run venue: the same operator controls task dispatch, execution, and evaluation.”

Checking the files later does not change who ran it.

Why believe it

You do not have to. Recompute it.

  1. Download the bundle (98.1 MB) and unpack it. It unpacks to one folder named by its evidence ID; rename that folder to bundle.
  2. Run the checker the bundle names on that folder:
    npx @colophon-claims/verify@0.2.1 ./bundle

Free, for anyone, on the published files. npx fetches the checker from npm; the check itself opens no network connection, needs no account and uploads nothing. Checking ends at the math, not at honesty.

Evidence

Evidence ID
de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774
Format
benchmark-product-public-bundle/7
Size
185.4 MB, 46,017 files
Download
tar.gz archive, 98.1 MB

The evidence ID is the SHA-256 of bundle.json, the bundle's file list, which records the SHA-256 of every other file.

Other IDs, the timestamp proof and every file

Check it step by step

curl -fLO 'https://github.com/colophon-claims/locomo-judge-report/releases/download/locomo-judge-report-run-completion-2026-09-01/de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774.tar.gz'
tar -xzf 'de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774.tar.gz'
mv 'de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774' bundle
npx @colophon-claims/verify@0.2.1 ./bundle

Needs Node 22 or newer. The checker reports 7 checks: manifest, evidence-closure, trust, matrix-rederivation, report-verification, claim-consistency, integrity-anchors. Releases before 0.2.1 refuse this format, 0.2.0 included, so run the exact version above.

Other IDs

Board key: the locked methodbenchmark.json
9ae50617f9112b750518c04309b96648207f6d0e17ba044a077d0d5185b84c9e

The SHA-256 of the method as locked. It is no official suite, so claims sealed on this same method share this key.

Run recordrun.json
481c9680e005eb7814f7120a56ef2242ba8167802571416ec363f6a743a6730e

Unique to this claim. The timestamp proof covers it. Its close time, 2026-08-29 16:30:51 UTC, is the sealed time printed above.

Reportreport.json
7869bab3a30defb29f171d7f572e68020edca789cfa7695ebe0a4621a9be5edd
Signed reportreport-envelope.json
4b16846a5ca64585867c55ce853afc0cdaafaaa947de7b6c15cb4de130396fe1
Result matrixmatrix.json
04de526b99dfa7180c875108101eaacbdfabe9a50a90ef317eca4c879ce1eb35
Reading record for this page
67feb4bce67084c1f53bc9c9110e640e9826c322bace8040742faece9c3e7798

Supplied at publication beside the bundle, not inside it, so the bundle stays byte for byte as its run produced it.

Timestamp proof

Timestamp proofproof file 18a34061…d54d
18a34061bb1d7dce5ef8eab9e69371e9db4318909d065a955154f5115190d54d

An OpenTimestamps proof, covering the run record, which the claim names as its lock. State, as the proof records it: pending. The proof records a submission and no confirmation yet, so it dates nothing yet.

A timestamp dates the bytes it covers; it does not validate the method, the run, or the conclusions.

Every file

The bundle holds 41,678 evidence records, 1 timestamp proof, 4,320 raw grading logs and the complete result matrix. Its file list, bundle.json, names 46,016 files, 177,261,232 bytes; with the list itself, 46,017 files, 185,417,411 bytes. Every file is served here byte for byte, under /reports/locomo-judge-report/bundle/. The main files:

FileBytesSHA-256
bundle.json8,156,179de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774
README.md9,577,757622a99456e7146b9048643970a435536763ac773caa5cc127a3d61d8d0d26c6a
badge.svg1,034446436ba65c1351cffaf79119e4502c23cd42e1cdb20b7d36379a598b411acad
benchmark.json24,4579ae50617f9112b750518c04309b96648207f6d0e17ba044a077d0d5185b84c9e
claim-package.json2,373,9793d6c18eaec73ee461fc6332c59ec4286eabfe55c4185e1eac92fe8bfe957245f
evidence.json4,414,042e0e0947a6ce77eacd3c9c0c234a7dafafe9e01accaec08b089be042b77516da0
index.html12,407,8330e9bdf0551704b0f4e64c42fb5af8ec72b2c91e011b1cb9d48b259fb13aa0adb
matrix.json4,092,24404de526b99dfa7180c875108101eaacbdfabe9a50a90ef317eca4c879ce1eb35
qualification.json316,105b1429f4835a1b3a351955ee3e9c166e7eae319ccac64968ee41c530694fc9670
report-envelope.json1,581,9084b16846a5ca64585867c55ce853afc0cdaafaaa947de7b6c15cb4de130396fe1
report.json1,186,2367869bab3a30defb29f171d7f572e68020edca789cfa7695ebe0a4621a9be5edd
run.json3,879481c9680e005eb7814f7120a56ef2242ba8167802571416ec363f6a743a6730e
share.txt270b8096b7bf29aa8949605146ce92a0d7bbd06562d5aade40f8a9e249ac1176774
social-card.svg1,227e22ef55d432e0462c02228d572c513d4350df566048d48ae620e8b560d2ec915
static-bundle.json2240a966bd90be1c1d602784753abc91187357ea5b915276bf9717af34c1bf748bc
trust/public-keys.json909a08be366afb0bd3191335b57bac4fba044948a480b73274818bb2654734ff691
verdicts.json1,242,0616e3fa20e84fb037dcff4652658b84b900a6115806093571e441852344bbcc775
verification/assembly.jsonl15,860,3134479d68e9191f0b867825a15516b027ef027375cd0bc150b13fb5ab8141ab458

Cite this claim

Judging the LoCoMo judges. Claim by signing key did:key:z6Mkmz8SWiUshwDMszSqwzMngt7ScNmsEj7vZoFRLjpcZTm9, sealed 2026-08-29, self-run. benchmark-product-public-bundle/7, evidence ID de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774. Report 7869bab3a30defb29f171d7f572e68020edca789cfa7695ebe0a4621a9be5edd. LoCoMo judge report frozen bank: 240 items, 6 grading prompts, 3 judge calls per item and prompt. https://colophon.claims/reports/locomo-judge-report/ Check: npx @colophon-claims/verify@0.2.1 ./bundle

The claimant's report

Below is the claimant's report. Its prose is taken from the claimant's reading record; the headings, charts, tables and the short notes around them are Colophon's, drawn from the sealed numbers. The lead above is Colophon's.

Wording revised . The sealed evidence bundle and every number are byte-identical and unchanged.

About this benchmark

LoCoMo is a benchmark for long-term conversational memory. Its published scores are often produced with an LLM judge: a model reads a question, a reference answer, and a candidate answer, then decides whether the candidate should count as correct.

The problem is that there is no single LoCoMo LLM judge. Different published results use different judge prompts, models, and inputs, then place the resulting scores next to one another as though the grading instrument had stayed fixed.

This benchmark measures how much that choice matters. Six grading configurations judged the same 240 answers on the same model snapshot. The configurations varied the judge prompt; one paired configuration also varied whether the judge received source evidence. Two additional tests examined consistency and behavior when the official answer key is wrong.

The central finding is simple: changing only the grader moved agreement with the same correctness labels from 60.8% to 87.9%.

The spread caused by changing the grader was 27.1 percentage points on the same answers. That is larger than many of the differences used to compare memory systems. Two scores produced by different graders are measurements from different instruments. Their numerical difference cannot be attributed to the memory systems alone.

Every grader was much more likely to accept a vague, on-topic wrong answer than a specific wrong answer. Three of five graders were internally inconsistent on list-answer cases involving subsets and supersets. When the answer key was wrong, graders followed the broken key against a true answer between 10% and 70% of the time.

This is a benchmark of graders, not memory systems. It does not rank or re-score any memory product, and it cannot establish which published system is better.

The five questions and their answers

These questions were published in the experiment design before the benchmark ran.

  1. Q1

    How often does each judge accept a known-wrong answer, and does the type of wrong answer matter?

    A great deal. Acceptance ranged from 2.5% to 28.8% for specific wrong answers and from 32.5% to 88.8% for vague, on-topic wrong answers. Every judge was substantially more forgiving of vagueness.

  2. Q2

    What does a stricter judge cost in rejected right answers?

    Very little when the reference answer is correct. Across 479 scored right answers, there was one rejection. The tradeoff appears when the reference answer itself is wrong.

  3. Q3

    Does changing the grader change the LoCoMo score?

    Yes. On identical answers, agreement with the labels ranged from 60.8% to 87.9% depending only on the grader.

  4. Q4

    Does showing the judge the source evidence matter?

    Yes. In this setup, evidence made the judge more permissive, not more accurate. The evidence-fed configuration accepted 7.7 percentage points more answers overall, with nearly all of the increase coming from vague wrong answers. Its overall agreement was lower, not higher.

  5. Q5

    When the answer key is wrong, does the judge follow the key or the truth?

    It depends on the grader. Against a true answer, broken-key following ranged from 10% to 70%. Every grader accepted all 20 true answers once the key was corrected.

A separate consistency test found that three of five applicable graders did not treat equivalent list-answer cases consistently.

1The design was posted publicly on 2026-08-18, before any official result existed.

Results

Agreement with the same correctness labels. Each answer was graded three times under every configuration; the final verdict was the majority of those calls. Intervals are 95% Wilson intervals.

Agreement with correctness labels

Same 240-item bank and model snapshot. The evidence-fed configuration has 233 scored items after seven exclusions. 95% Wilson intervals.

Strict-dial87.9%
Audited82.1%
Mem078.3%
Mem0 + evidence70.0%
Backboard68.3%
Revised60.8%
Changing only the grading configuration produced a 27.1-point spread on identical inputs.

Known-wrong answers accepted

Majority verdicts. Each class has 80 items; the evidence-fed arm scored 76 specific and 78 vague after declared exclusions.

Specific wrongVague wrong
Strict-dial2.5%32.5%
Audited11.3%42.5%
Mem012.5%52.5%
Mem0 + evidence14.5%75.6%
Backboard17.5%77.5%
Revised28.7%88.8%
Every judge was substantially more forgiving of answers that stayed on topic while avoiding the requested fact.
View exact rates, counts, and intervals
JudgeAgreement95% intervalAccepts specific-wrongAccepts vague-topical-wrongRejects correct
audited82.1% (197/240)76.7% to 86.4%11.3% (9/80)42.5% (34/80)0.0% (0/80)
backboard68.3% (164/240)62.2% to 73.9%17.5% (14/80)77.5% (62/80)0.0% (0/80)
mem078.3% (188/240)72.7% to 83.1%12.5% (10/80)52.5% (42/80)0.0% (0/80)
mem0-evidence70.0% (163/233)63.8% to 75.5%14.5% (11/76)75.6% (59/78)0.0% (0/79)
revised60.8% (146/240)54.5% to 66.8%28.7% (23/80)88.8% (71/80)0.0% (0/80)
strict-dial87.9% (211/240)83.2% to 91.5%2.5% (2/80)32.5% (26/80)1.3% (1/80)
2Spread runs from revised to strict-dial, 27.1 points apart on the same items.

Results from two additional tests

The wrong-key test directly answers Q5 and explains the exception in Q2. The consistency test checks a separate property: whether equivalent answers receive the same verdict.

Answers Q5 and qualifies Q2

When the answer key is wrong

Depending on the grader, 10% to 70% of true answers were rejected when the key was wrong. Every grader accepted all 20 once the key was corrected. This is where stricter grading creates a real cost.

Separate consistency diagnostic

Equivalent answers should receive the same verdict

Three of five graders changed their decision across equivalent list-answer cases. The two graders that stayed consistent accepted every probe, so consistency alone is not correctness.

This is a benchmark of graders, not memory systems. It does not rank or re-score any memory product, and it cannot establish which published system is better.

Recommendations

Six actions for more reliable LoCoMo judging

  1. Audit answer keys before tightening the judge. Strict grading caused only one false rejection across 479 right answers when the key was correct, but the strictest grader rejected 70% of true answers paired with broken keys. Adopt the audit's corrected keys, verify them, then standardize a stricter judge.
  2. Require answers to state the requested fact. Every judge struggled more with vague, on-topic answers; even the strictest accepted 32.5%. Test a judge that first identifies the exact value an answer gives, then compares it with the reference. If no value is given, the answer fails.
  3. Add a third verdict for a likely broken key. Every grader accepted all 20 true answers once the key was corrected. When an answer contradicts the reference but is supported by evidence, send it to review instead of automatically marking it wrong.
  4. Test consistency with both accept and reject cases. The two graders that passed the current test accepted every probe, so it rewarded leniency as well as consistency. A replacement should include cases where rejection is the only consistent decision.
  5. Spend evaluation effort on more items and better keys. Repeat disagreement was 1.6% overall and 2.9% at worst, while grader choice moved agreement by 27.1 points. On this bank, broader coverage and key auditing are more valuable than additional repeated calls.
  6. Tell the judge how to use evidence. Simply adding source evidence raised acceptance by 7.7 points and lowered agreement from 78.3% to 70.0%; almost all extra acceptances were wrong. A reference prompt should require explicit comparison with the evidence, not merely include it as context.

The practical recommendation is one versioned LoCoMo reference judge, used with audited answer keys.

It should identify the fact an answer commits to, compare that fact with the reference, and allow review when the key appears wrong. These design changes remain proposals for a follow-up benchmark, not findings of this run.

How to use LoCoMo scores today

The disclosure standard this benchmark supports

A LoCoMo score is not fully specified by the memory system's name and final percentage. It is the product of an answer pipeline and a grading pipeline. A comparable result must disclose the six choices that define those pipelines.

The six variables

VariableWhat it describes
Ingestion modelWhat reads the source conversation and builds the memory or index
Retrieval configurationHow information is selected, including search strategy, depth, and limits
Answer modelWhat produces the candidate answer
Answer promptThe instructions under which the candidate answer is produced
Judge modelWhat grades the candidate against the reference answer
Judge promptThe grading instructions and the complete input shown to the judge

Whether the judge receives source evidence belongs in the judge-prompt entry because it changes the judge's input. Parser behavior and repeated-call aggregation should be reported alongside the judge configuration.

Three ways to report each variable

Each of the six variables must have one of three states:

State shown in a reportMeaning
Measured in this benchmarkThe experiment fixed and recorded the variable itself
Reported by the publisherThe publisher stated the choice, but this experiment did not measure or independently establish it
Not reportedThe choice is unknown, not stated, or not applicable

Unknown information must be reported as not reported, not omitted or inferred from a repository name, model family, or surrounding context. Evidence belongs only with a choice measured in the benchmark. A publisher statement may be cited, but it should not be presented as an independent measurement.

Why the benchmark supports this standard

This benchmark directly measures the importance of the grading variables. Holding the answer bank and judge model fixed while changing the judge prompt moved agreement from 60.8% to 87.9%. Holding the Mem0 prompt family fixed while adding source evidence changed acceptance by 7.7 percentage points. The judge prompt and the judge's input therefore change the meaning of the resulting score.

The experiment did not vary the first four variables. They still need disclosure because they determine which candidate answers reach the grader. Without them, a reader cannot distinguish a memory-system effect from a retrieval, generation, or answer-style effect. The benchmark establishes the grader side of the problem; the complete reporting standard covers the whole measurement pipeline.

A LoCoMo score without all six entries is an incompletely described measurement. Scores produced under different disclosed configurations should not be treated as directly comparable.

This report's disclosure

VariableStateDeclaration
Ingestion modelNot reportedThe published source materials do not identify what produced the upstream memories or indexes
Retrieval configurationNot reportedThe published source materials do not identify the retrieval settings used upstream
Answer modelNot reportedThe candidate-answer files do not identify the model that wrote every answer
Answer promptNot reportedThe candidate-answer files do not identify the complete answering instructions
Judge modelMeasured in this benchmarkAll grading calls used gpt-4o-mini-2024-07-18 at temperature zero
Judge promptMeasured in this benchmarkThe six grading configurations and their input shapes were fixed and recorded by this benchmark

Four entries are marked not reported. That is the honest description of a grader benchmark that used candidate answers produced elsewhere. This report makes no claim about upstream choices that its source materials cannot establish.

How the benchmark was run

Six grading configurations judged the same 240 answers using the same model version and settings.

Grading configurationPrompt tested
auditedThe EverMemOS-derived prompt reconstructed in the community audit
backboardBackboard's posted prompt
mem0Mem0's posted prompt
mem0-evidenceThe Mem0 prompt with the dataset's source evidence added
revisedThe successor rubric from mem0ai/memory-benchmarks
strict-dialThe stricter community variant posted by @dial481
Held constant
Judge model
gpt-4o-mini-2024-07-18
Model settings
gpt-4o-mini-2024-07-18, temperature 0.0
Evaluation software
inspect-ai-judge@1
Repeated grading
Each answer was graded 3 times per configuration
Final verdict
Majority of the 3 calls
Confidence intervals
95% Wilson intervals
Answers and labels
Answer bank
240 candidate answers, balanced across three correctness classes
Correctness labels
A model screened every label; the claimant reviewed flagged cases and a random sample
Grading prompts
6, already present in the LoCoMo ecosystem or community discussion
Run operator
The claimant, signing key z6Mkmz8SWi…pcZTm9

How the additional tests were run

The consistency test used 12 constructed list-answer probes, five applicable grading configurations, and three calls per probe, for 180 calls. The wrong-key test used 20 questions under both the official and corrected answer keys, six configurations, and three calls per case, for 720 calls. All planned calls completed.

What this benchmark does not establish

Data and verification

Download the full result data, grading outputs, prompts, and analysis files from the published evidence package. Colophon packages these materials and checks that the files used by this report have not changed since publication.

File verification confirms the integrity of the published materials. It does not prove that the correctness labels are right or independently verify the model provider's internal execution. Independent timestamping is pending.

Technical verification details
Companion runRecordSHA-256
Consistency gateRun06ae9e458b12562015296228297828e22f291ec0c73ec6833eec4288f0511be6
Consistency gateResult matrixaaeddc612a8625d598ccce4ffa03318ed546cce3854912244baa4cce91a79d60
Consistency gateCurrent bundled1535f32bfa850f2e5ecedcae4be97e17e6dedff37aebe412f2ef6d36fc6a404
Corrupt-key testRunbdc3b0e2ac4c22e9a859b5879118af7df3939ad4501ea3afd4c8ede0770e509d
Corrupt-key testResult matrix2a788dc9cdcaa53260dac8b0f7728d3bfd62aac24ccfcbaddc213fa0729ea2d7
Corrupt-key testCurrent bundle271f87db4e616992dea75483ec69e92624ab7fc84aeb32e6a0bc82674dc506ed

Each bundle was verified from outside the workspace that produced it, and every tested tamper variant was rejected.

The bundle, its IDs and the check command are in the evidence row at the top of this page.

Materials, credit, and license

This benchmark builds on the published materials in dial481/locomo-audit, with the author's stated permission, and on judge prompts posted by Backboard, Mem0, and community contributors in the LoCoMo discussion. The audited judge prompt derives from EverMemOS, and the revised rubric from mem0ai/memory-benchmarks, both under Apache-2.0. The consistency examples originated with @AnitaLeungxx, and the strict variant and audit materials with @dial481.

The underlying dataset is LoCoMo. LoCoMo and the audit annotations are CC BY-NC 4.0. This report and the selected derived materials are released under the same non-commercial license, with source pointers rather than full conversation data.

Every planned call, accounted for

Every judge call the run planned is counted here, including the ones that failed. No flag drops a row.

All 4,320 planned judge calls completed. Twenty-two responses could not be parsed, which excluded seven evidence-fed items from that comparison.

4320 expected judge calls
completed 4,320 / 4,320lost 0 / 4,320
Planned judge calls, by outcome
OutcomeCalls
Judged, with a verdict the parser could read4,298
Judged, but the response could not be parsed22
Not judged: lost or not run0
All planned calls4,320
Items excluded, by grading configuration
ConfigurationItems
audited0
backboard0
mem00
mem0-evidence7
revised0
strict-dial0
Total7

The record's reason for the exclusions: “Unparseable responses were recorded as neutral rather than converted into rejections, all of them in the evidence-fed configuration. The evidence-fed items consequently left without a valid majority appear as explicit exclusions.”

The method's pre-set minimum for a complete run: 99.5% of planned calls judged. This run judged 4,320 of 4,320; its recorded outcome is complete.

Repeat stability. Repeat disagreement was 1.6% overall and 2.9% in the least stable grading configuration.

3Responses that could not be parsed were treated as neither correct nor incorrect. All 22 occurred in the evidence-fed configuration.