All posts
Research

Five Ways AI Memory Fails, and the Design That Answers Each

Published  · 5 min read

By Alexander Snyder, inventor of Edwin

AI memory goes wrong in five predictable ways, from trusting a confident guess to losing track of what changed. Here is each one, why it costs you, and how Edwin is built around it.

AI MemoryDesignEvidence
Short answer: AI memory fails in five predictable ways: it trusts a confident guess, mistakes an update for an error, cries wolf about contradictions, costs too much to ask, and locks your history inside one product. Edwin is built around all five, with honest labels for what's been checked, dated facts, contradictions that count only when two AI providers agree or you confirm them, answers drawn from a small slice of your history, and a memory file you own.

Edwin is a memory layer for the AI tools you already use. It works with any MCP-compatible client, such as Claude, and remembers what you've told your AI: when it was true, what was decided, and what's actually been checked. These five failures shaped its design. Where we've measured the answer, the result sits right beside it, with the full record on our research page.

Why isn't a confident answer a correct one?

Because confidence measures how sure a model sounds, not whether anyone has checked the claim. A memory that stores whatever sounds confident will eventually hand your AI something wrong, stated with total conviction. Edwin keeps those two questions apart and labels every memory with what has actually happened to it.

So beside every memory, Edwin shows one of four plain labels. A second AI provider challenges stored facts in the background, and the label says how each one has fared:

  • Survived challenge: a second AI provider tried to dispute it and couldn't.
  • Not confirmed: it was challenged and the support fell short.
  • Disputed: the providers disagree about it.
  • Not yet challenged: no second provider has checked it yet.
The result: your AI can tell the memories that have held up from the ones nobody has tested yet.Evidence: checking facts and showing the AI those labels raised accuracy 5.8 points on STALE, a third-party benchmark built on facts that change. (pre-registered, Oct 2026)

When is a change not a contradiction?

When one fact was true then and another is true now. If your budget went from $500K to $750K, that's an update, not an error. Edwin stores every fact with the date it was true, so it records the change as an update and keeps the old value on record with its date.

Budgets, owners and plans change all the time. A memory that ignores time either flags every change as a conflict or overwrites the old value and loses the history. Neither helps when you need to know what's true today and what it used to be.

Edwin's answer: every fact carries the date it was true.Evidence: with source dates, Edwin classified 50 of 50 updates correctly; without them, 23 of 50. (internal test, Feb 2026)

Decisions get the same care. Edwin keeps the decisions from your meetings as top-priority memories, tied to the meeting they came from and the facts discussed with them, so they don't get buried under newer conversation.

How do you catch contradictions without crying wolf?

Make every flag earn its place. A checker that flags too much gets ignored, and then it protects no one. In Edwin, a contradiction counts only when two different AI providers agree or you confirm it, so one model's hunch never sets off the alarm on its own.

Edwin's answer: build the detector for precision first, then add a second opinion.Evidence: on its own, Edwin's contradiction detector raised zero false alarms on a 150-pair test while catching 68 of 100 real contradictions. A second-provider check now sits on top of it. (internal test, Apr 2026)

Why should memory make each question cheaper?

The simplest memory is pasting your whole history into every question. It works while the history is short. As it grows, every question carries everything you've ever said, and you pay for all of it, every time.

Edwin's answer: answer from the few facts that matter.Evidence: Edwin answered from about 0.3% of the tokens of pasting the whole history in, at the same overall accuracy: about 125× cheaper per question. (Aug 2026, self-run)

Who owns your AI's memory?

Memory built inside one AI app lives on that company's servers, under its terms. As your AI learns more about your work, that history becomes more valuable, and more worth owning.

Edwin's answer: your memory is a SQLite file on your own Mac, with encrypted nightly backups to a location you control (iCloud Drive by default). It works with any MCP-compatible AI client, so changing tools doesn't mean starting over.

Local-first, not local-only: AI model providers extract and check facts. Our privacy policy lists what goes where.

+5.8 pts
accuracy from checking facts and showing labels · pre-registered, Oct 2026
50 of 50
updates classified correctly with source dates (23 without) · Feb 2026
0
false alarms from the contradiction detector on its own, catching 68 of 100 · Apr 2026
~0.3%
of the tokens of pasting your whole history in · Aug 2026

For technical readers: where do these numbers come from?

From our own measurement record, where every result carries its date and sample size. The October 2026 results were pre-registered: we wrote down the test and the success bar before running it. STALE is a third-party benchmark, and we ran it ourselves. The harness and spend ledger are available on request. The record publishes the experiments that didn't work alongside the ones that did, in the October 2026 measurements and the experiment log. Edwin is patent pending; the provisional application was filed February 26, 2026.

What does this add up to?

A memory your AI can rely on. It says what's been checked, it knows when each fact was true, it raises a contradiction only when two providers agree or you confirm it, it answers from a small slice of your history, and it belongs to you. Over the coming weeks, each of these gets a post of its own in this series.

Edwin is in private development.

Edwin remembers what you've told your AI: when it was true, what was decided, and what's actually been checked. Send us a note and we'll tell you when it opens.