All posts
Research

How to Read an AI Memory Benchmark: A Buyer's Guide

Published · Updated 6 min read
By Alexander Snyder, inventor of Edwin

Published AI memory scores often move when someone else re-runs them. Here is what to look for in a benchmark claim, the questions to ask any vendor, and what the public record shows, dated and linked.

BenchmarksBuyer's GuideAI Memory
Short answer: A benchmark score is only as meaningful as the setup behind it. Before trusting one, ask who ran it, with which harness, judge and answering model, and whether anyone else has reproduced it. Published AI memory scores often move when re-run: Zep, for example, corrected its own LoCoMo figure from about 84 to 75.14 in May 2025 after a competitor re-scored it.

Why do memory scores change when someone else runs them?

A memory benchmark score depends on more than the memory system. The harness that feeds conversations in, the model that writes the answers, the judge that grades them and the prompts each one sees all move the number. Change any of them and the score changes, sometimes by a lot. Three dated examples from the public record (sources checked October 3, 2026):

None of these re-runs is the last word. Each uses its own harness, and two come from a competitor with its own interest. The lesson is the pattern: a single number, from anyone, means most when you can see how it was produced.

81.08
Mem0 on LoCoMo, independent harness (published 92.5) · arXiv:2601.07978 v3, May 2026
93.4 → 73.8
Mem0 LongMemEval · maximem.ai, a rival vendor, June 2026
~84 → 75.14
Zep LoCoMo · Zep's own correction, May 12, 2025
~0.3%
Edwin's tokens vs full context, same accuracy · STALE, Aug 2026, self-run

What should you ask an AI memory vendor?

  • Who ran it? A vendor's own run, an independent lab and a competitor's re-run each tell you something different.
  • Has anyone reproduced it? A score that holds up on someone else's harness is worth more than a higher one that hasn't been re-run.
  • Which harness, judge and answering model? A stronger answering model can lift any memory system's score, so a comparison means something only when these match.
  • Was the test planned in advance? Pre-registration (writing down the test and the success bar before running it) and confidence intervals separate a real gain from noise.
  • What does each answer cost? Tokens per question tell you whether a score is affordable at your volume.
  • Where does it lose? Ask for results by question type. They show where a system is weaker, and a vendor who shares them is one you can plan around.

What do the main AI memory systems publish? (checked October 3, 2026)

Mem0

Available open-source and hosted. It publishes 92.5 on LoCoMo and 94.4 on LongMemEval for its 2026 algorithm, at under 7,000 tokens per retrieval call (Mem0 blog (opens in a new tab)). An independent run of the open-source v2 on another harness scored 81.08 on LoCoMo; the LongMemEval re-run above was of its earlier 93.4.

Zep and Graphiti

A bi-temporal knowledge graph that records when facts were introduced and when they changed; Zep's paper (arXiv:2501.13956 (opens in a new tab), January 2025) describes its bi-temporal edge invalidation. Graphiti, its graph engine, is open source and can be self-hosted.

Hindsight

The paper is "Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects" (opens in a new tab) (arXiv:2512.12818 (opens in a new tab), December 2025). It reports 91.4% on LongMemEval and 89.61% on LoCoMo with a scaled backbone, and 83.6% on LongMemEval with an open-source 20B model. The code (opens in a new tab) is MIT-licensed and self-hostable (Docker or pip), with a managed option.

Cloudflare Agent Memory

Announced April 17, 2026 (opens in a new tab). A verifier runs eight checks on a memory, among them temporal accuracy and whether an inferred fact is supported by the conversation, and a newer memory on the same topic supersedes the older one through a version chain.

ChatGPT, Claude and Gemini memory

These store what the user says and retrieve it later, on each company's servers, under each company's retention and training settings.

How does Edwin report its own numbers?

We hold our results to the same questions. Here is our headline cost result, reported the way we think every memory benchmark should be:

  • The claim: Edwin answered from about 0.3% of the tokens of pasting the whole history in, at the same overall accuracy: 23.7% against 22.8% for full context. That makes it about 125× cheaper per question, and its setup pays for itself after about 3 questions.
  • The setup: STALE, a deliberately hard benchmark in which every question is about a fact that changed; 180 pre-registered held-out scenarios, with both arms answered by the same model and graded by the same judge.
  • Who ran it: we did, in August 2026. The harness, judge prompts and spend ledger are available on request.
  • The rest of the record: results by question type, later runs and the experiments that didn't work are in our research record.

Publishing the whole record is a deliberate choice. Memory you rely on should be measured in public, and the more of the field that does the same, the easier a buyer's job becomes. For what Edwin is built to get right, see five ways AI memory fails.

Edwin is in private development.

Edwin remembers what you've told your AI: when it was true, why you decided it, and what's actually been checked. Send us a note and we'll tell you when it opens.