The claimant chose them.Not a fixed public set, not drawn by lottery. Questions sampled by the claimant's seeded script; a model screened a 664-item pool, the operator hand-reviewed 255, and 137 were excluded and replaced to keep the 80/80/80 balance.
Method
LoCoMo judge report frozen bank 1.0.0, written by the claimant. All 240 planned items, 4,320 planned calls.
Arms
Six grading prompts, supplied by the claimant. Colophon does not look inside them.
Scorer
The method's own, fixed in the sealed run: gpt-4o-mini-2024-07-18 at temperature 0.0, through inspect-ai-judge 1. The majority of 3 calls decides each answer.
Who ran it
Self-run
The claimant ran it. The sealed disclosure opens:
“This is a local, self-run venue: the same operator controls task dispatch, execution, and evaluation.”
Checking the files later does not change who ran it.
Why believe it
You do not have to. Recompute it.
Download the bundle (98.1 MB) and unpack it. It unpacks to one folder named by its evidence ID; rename that folder to bundle.
Run the checker the bundle names on that folder:
npx @colophon-claims/verify@0.2.1 ./bundle
Free, for anyone, on the published files. npx fetches the checker from npm; the check itself opens no network connection, needs no account and uploads nothing. Checking ends at the math, not at honesty.
Needs Node 22 or newer. The checker reports 7 checks: manifest, evidence-closure, trust, matrix-rederivation, report-verification, claim-consistency, integrity-anchors. Releases before 0.2.1 refuse this format, 0.2.0 included, so run the exact version above.
An OpenTimestamps proof, covering the run record, which the claim names as its lock. State, as the proof records it: pending. The proof records a submission and no confirmation yet, so it dates nothing yet.
A timestamp dates the bytes it covers; it does not validate the method, the run, or the conclusions.
Every file
The bundle holds 41,678 evidence records, 1 timestamp proof, 4,320 raw grading logs and the complete result matrix. Its file list, bundle.json, names 46,016 files, 177,261,232 bytes; with the list itself, 46,017 files, 185,417,411 bytes. Every file is served here byte for byte, under /reports/locomo-judge-report/bundle/. The main files:
Judging the LoCoMo judges. Claim by signing key did:key:z6Mkmz8SWiUshwDMszSqwzMngt7ScNmsEj7vZoFRLjpcZTm9, sealed 2026-08-29, self-run. benchmark-product-public-bundle/7, evidence ID de169c04a24bbb4d9d5b52e398b8bfe92e939ba4d8a8c9e6e3e551cf10aa3774. Report 7869bab3a30defb29f171d7f572e68020edca789cfa7695ebe0a4621a9be5edd. LoCoMo judge report frozen bank: 240 items, 6 grading prompts, 3 judge calls per item and prompt. https://colophon.claims/reports/locomo-judge-report/ Check: npx @colophon-claims/verify@0.2.1 ./bundle
The claimant's report
Below is the claimant's report. Its prose is taken from the claimant's reading record; the headings, charts, tables and the short notes around them are Colophon's, drawn from the sealed numbers. The lead above is Colophon's.
Wording revised . The sealed evidence bundle and every number are byte-identical and unchanged.
About this benchmark
LoCoMo is a benchmark for long-term conversational memory. Its published scores are often produced with an LLM judge: a model reads a question, a reference answer, and a candidate answer, then decides whether the candidate should count as correct.
The problem is that there is no single LoCoMo LLM judge. Different published results use different judge prompts, models, and inputs, then place the resulting scores next to one another as though the grading instrument had stayed fixed.
This benchmark measures how much that choice matters. Six grading configurations judged the same 240 answers on the same model snapshot. The configurations varied the judge prompt; one paired configuration also varied whether the judge received source evidence. Two additional tests examined consistency and behavior when the official answer key is wrong.
The central finding is simple: changing only the grader moved agreement with the same correctness labels from 60.8% to 87.9%.
The spread caused by changing the grader was 27.1 percentage points on the same answers. That is larger than many of the differences used to compare memory systems. Two scores produced by different graders are measurements from different instruments. Their numerical difference cannot be attributed to the memory systems alone.
Every grader was much more likely to accept a vague, on-topic wrong answer than a specific wrong answer. Three of five graders were internally inconsistent on list-answer cases involving subsets and supersets. When the answer key was wrong, graders followed the broken key against a true answer between 10% and 70% of the time.
This is a benchmark of graders, not memory systems. It does not rank or re-score any memory product, and it cannot establish which published system is better.
The five questions and their answers
These questions were published in the experiment design before the benchmark ran.
Q1
How often does each judge accept a known-wrong answer, and does the type of wrong answer matter?
A great deal. Acceptance ranged from 2.5% to 28.8% for specific wrong answers and from 32.5% to 88.8% for vague, on-topic wrong answers. Every judge was substantially more forgiving of vagueness.
Q2
What does a stricter judge cost in rejected right answers?
Very little when the reference answer is correct. Across 479 scored right answers, there was one rejection. The tradeoff appears when the reference answer itself is wrong.
Q3
Does changing the grader change the LoCoMo score?
Yes. On identical answers, agreement with the labels ranged from 60.8% to 87.9% depending only on the grader.
Q4
Does showing the judge the source evidence matter?
Yes. In this setup, evidence made the judge more permissive, not more accurate. The evidence-fed configuration accepted 7.7 percentage points more answers overall, with nearly all of the increase coming from vague wrong answers. Its overall agreement was lower, not higher.
Q5
When the answer key is wrong, does the judge follow the key or the truth?
It depends on the grader. Against a true answer, broken-key following ranged from 10% to 70%. Every grader accepted all 20 true answers once the key was corrected.
A separate consistency test found that three of five applicable graders did not treat equivalent list-answer cases consistently.
Agreement with the same correctness labels. Each answer was graded three times under every configuration; the final verdict was the majority of those calls. Intervals are 95% Wilson intervals.
Agreement with correctness labels
Same 240-item bank and model snapshot. The evidence-fed configuration has 233 scored items after seven exclusions. 95% Wilson intervals.
0255075100%
Strict-dial87.9%
Audited82.1%
Mem078.3%
Mem0 + evidence70.0%
Backboard68.3%
Revised60.8%
Changing only the grading configuration produced a 27.1-point spread on identical inputs.
Known-wrong answers accepted
Majority verdicts. Each class has 80 items; the evidence-fed arm scored 76 specific and 78 vague after declared exclusions.
Specific wrongVague wrong
Strict-dial2.5%32.5%
Audited11.3%42.5%
Mem012.5%52.5%
Mem0 + evidence14.5%75.6%
Backboard17.5%77.5%
Revised28.7%88.8%
Every judge was substantially more forgiving of answers that stayed on topic while avoiding the requested fact.
The wrong-key test directly answers Q5 and explains the exception in Q2. The consistency test checks a separate property: whether equivalent answers receive the same verdict.
Answers Q5 and qualifies Q2
When the answer key is wrong
Depending on the grader, 10% to 70% of true answers were rejected when the key was wrong. Every grader accepted all 20 once the key was corrected. This is where stricter grading creates a real cost.
Separate consistency diagnostic
Equivalent answers should receive the same verdict
Three of five graders changed their decision across equivalent list-answer cases. The two graders that stayed consistent accepted every probe, so consistency alone is not correctness.
This is a benchmark of graders, not memory systems. It does not rank or re-score any memory product, and it cannot establish which published system is better.
Recommendations
Six actions for more reliable LoCoMo judging
Audit answer keys before tightening the judge. Strict grading caused only one false rejection across 479 right answers when the key was correct, but the strictest grader rejected 70% of true answers paired with broken keys. Adopt the audit's corrected keys, verify them, then standardize a stricter judge.
Require answers to state the requested fact. Every judge struggled more with vague, on-topic answers; even the strictest accepted 32.5%. Test a judge that first identifies the exact value an answer gives, then compares it with the reference. If no value is given, the answer fails.
Add a third verdict for a likely broken key. Every grader accepted all 20 true answers once the key was corrected. When an answer contradicts the reference but is supported by evidence, send it to review instead of automatically marking it wrong.
Test consistency with both accept and reject cases. The two graders that passed the current test accepted every probe, so it rewarded leniency as well as consistency. A replacement should include cases where rejection is the only consistent decision.
Spend evaluation effort on more items and better keys. Repeat disagreement was 1.6% overall and 2.9% at worst, while grader choice moved agreement by 27.1 points. On this bank, broader coverage and key auditing are more valuable than additional repeated calls.
Tell the judge how to use evidence. Simply adding source evidence raised acceptance by 7.7 points and lowered agreement from 78.3% to 70.0%; almost all extra acceptances were wrong. A reference prompt should require explicit comparison with the evidence, not merely include it as context.
The practical recommendation is one versioned LoCoMo reference judge, used with audited answer keys.
It should identify the fact an answer commits to, compare that fact with the reference, and allow review when the key appears wrong. These design changes remain proposals for a follow-up benchmark, not findings of this run.
How to use LoCoMo scores today
Compare scores only when the grading configuration matches; otherwise, do not treat the difference as a memory-system effect.
Require the judge model, prompt, evidence input, parser behavior, and repeated-call rule.
Publish error rates separately for specific wrong and vague wrong answers.
Treat small score gaps cautiously when grader choice can move agreement by 27.1 points.
The disclosure standard this benchmark supports
A LoCoMo score is not fully specified by the memory system's name and final percentage. It is the product of an answer pipeline and a grading pipeline. A comparable result must disclose the six choices that define those pipelines.
The six variables
Variable
What it describes
Ingestion model
What reads the source conversation and builds the memory or index
Retrieval configuration
How information is selected, including search strategy, depth, and limits
Answer model
What produces the candidate answer
Answer prompt
The instructions under which the candidate answer is produced
Judge model
What grades the candidate against the reference answer
Judge prompt
The grading instructions and the complete input shown to the judge
Whether the judge receives source evidence belongs in the judge-prompt entry because it changes the judge's input. Parser behavior and repeated-call aggregation should be reported alongside the judge configuration.
Three ways to report each variable
Each of the six variables must have one of three states:
State shown in a report
Meaning
Measured in this benchmark
The experiment fixed and recorded the variable itself
Reported by the publisher
The publisher stated the choice, but this experiment did not measure or independently establish it
Not reported
The choice is unknown, not stated, or not applicable
Unknown information must be reported as not reported, not omitted or inferred from a repository name, model family, or surrounding context. Evidence belongs only with a choice measured in the benchmark. A publisher statement may be cited, but it should not be presented as an independent measurement.
Why the benchmark supports this standard
This benchmark directly measures the importance of the grading variables. Holding the answer bank and judge model fixed while changing the judge prompt moved agreement from 60.8% to 87.9%. Holding the Mem0 prompt family fixed while adding source evidence changed acceptance by 7.7 percentage points. The judge prompt and the judge's input therefore change the meaning of the resulting score.
The experiment did not vary the first four variables. They still need disclosure because they determine which candidate answers reach the grader. Without them, a reader cannot distinguish a memory-system effect from a retrieval, generation, or answer-style effect. The benchmark establishes the grader side of the problem; the complete reporting standard covers the whole measurement pipeline.
A LoCoMo score without all six entries is an incompletely described measurement. Scores produced under different disclosed configurations should not be treated as directly comparable.
This report's disclosure
Variable
State
Declaration
Ingestion model
Not reported
The published source materials do not identify what produced the upstream memories or indexes
Retrieval configuration
Not reported
The published source materials do not identify the retrieval settings used upstream
Answer model
Not reported
The candidate-answer files do not identify the model that wrote every answer
Answer prompt
Not reported
The candidate-answer files do not identify the complete answering instructions
Judge model
Measured in this benchmark
All grading calls used gpt-4o-mini-2024-07-18 at temperature zero
Judge prompt
Measured in this benchmark
The six grading configurations and their input shapes were fixed and recorded by this benchmark
Four entries are marked not reported. That is the honest description of a grader benchmark that used candidate answers produced elsewhere. This report makes no claim about upstream choices that its source materials cannot establish.
How the benchmark was run
Six grading configurations judged the same 240 answers using the same model version and settings.
Grading configuration
Prompt tested
audited
The EverMemOS-derived prompt reconstructed in the community audit
backboard
Backboard's posted prompt
mem0
Mem0's posted prompt
mem0-evidence
The Mem0 prompt with the dataset's source evidence added
revised
The successor rubric from mem0ai/memory-benchmarks
strict-dial
The stricter community variant posted by @dial481
Held constant
Judge model
gpt-4o-mini-2024-07-18
Model settings
gpt-4o-mini-2024-07-18, temperature 0.0
Evaluation software
inspect-ai-judge@1
Repeated grading
Each answer was graded 3 times per configuration
Final verdict
Majority of the 3 calls
Confidence intervals
95% Wilson intervals
Answers and labels
Answer bank
240 candidate answers, balanced across three correctness classes
Correctness labels
A model screened every label; the claimant reviewed flagged cases and a random sample
Grading prompts
6, already present in the LoCoMo ecosystem or community discussion
Run operator
The claimant, signing key z6Mkmz8SWi…pcZTm9
How the additional tests were run
The consistency test used 12 constructed list-answer probes, five applicable grading configurations, and three calls per probe, for 180 calls. The wrong-key test used 20 questions under both the official and corrected answer keys, six configurations, and three calls per case, for 720 calls. All planned calls completed.
What this benchmark does not establish
It does not evaluate or rank memory systems.
It tests one dated judge-model snapshot on a balanced 240-answer diagnostic bank. Rates may differ on another model, snapshot, or answer distribution.
The correctness labels were model-screened and selectively reviewed by the authors, not independently labelled by multiple human annotators.
The 20 broken-key items and 12 consistency probes demonstrate mechanisms; they do not estimate how common those problems are in ordinary benchmark runs.
The authors operated the experiment themselves. Colophon verifies the published files and their internal consistency, but not the provider's internal execution or the authors' organizational independence.
Data and verification
Download the full result data, grading outputs, prompts, and analysis files from the published evidence package. Colophon packages these materials and checks that the files used by this report have not changed since publication.
File verification confirms the integrity of the published materials. It does not prove that the correctness labels are right or independently verify the model provider's internal execution. Independent timestamping is pending.
Each bundle was verified from outside the workspace that produced it, and every tested tamper variant was rejected.
The bundle, its IDs and the check command are in the evidence row at the top of this page.
Materials, credit, and license
This benchmark builds on the published materials in dial481/locomo-audit, with the author's stated permission, and on judge prompts posted by Backboard, Mem0, and community contributors in the LoCoMo discussion. The audited judge prompt derives from EverMemOS, and the revised rubric from mem0ai/memory-benchmarks, both under Apache-2.0. The consistency examples originated with @AnitaLeungxx, and the strict variant and audit materials with @dial481.
The underlying dataset is LoCoMo. LoCoMo and the audit annotations are CC BY-NC 4.0. This report and the selected derived materials are released under the same non-commercial license, with source pointers rather than full conversation data.
Every planned call, accounted for
Every judge call the run planned is counted here, including the ones that failed. No flag drops a row.
All 4,320 planned judge calls completed. Twenty-two responses could not be parsed, which excluded seven evidence-fed items from that comparison.
4320 expected judge calls
completed 4,320 / 4,320lost 0 / 4,320
Planned judge calls, by outcome
Outcome
Calls
Judged, with a verdict the parser could read
4,298
Judged, but the response could not be parsed
22
Not judged: lost or not run
0
All planned calls
4,320
Items excluded, by grading configuration
Configuration
Items
audited
0
backboard
0
mem0
0
mem0-evidence
7
revised
0
strict-dial
0
Total
7
The record's reason for the exclusions: “Unparseable responses were recorded as neutral rather than converted into rejections, all of them in the evidence-fed configuration. The evidence-fed items consequently left without a valid majority appear as explicit exclusions.”
The method's pre-set minimum for a complete run: 99.5% of planned calls judged. This run judged 4,320 of 4,320; its recorded outcome is complete.
Repeat stability. Repeat disagreement was 1.6% overall and 2.9% in the least stable grading configuration.
3Responses that could not be parsed were treated as neither correct nor incorrect. All 22 occurred in the evidence-fed configuration.