All posts
Research

Five Ways AI Memory Fails, and How Edwin Answers Each

Published · Updated 6 min read
By Alexander Snyder, inventor of Edwin

AI memory goes wrong in five predictable ways, from trusting confident guesses to losing track of what changed. Here is each failure, why it matters, and the design choice Edwin makes to prevent it.

AI MemoryDesignResearch
Short answer: AI memory fails in five predictable ways: it treats a confident answer as a correct one, mistakes a change for a contradiction, misses real contradictions or raises false ones, lets its checkers drift into agreement, and keeps your history inside someone else's product. Edwin is designed around all five, with dated facts, honest labels, cross-provider checks and a memory file you own.

Edwin is a memory layer for the AI tools you already use. It works with any MCP-compatible client, such as Claude, and remembers what you've told your AI: when it was true, why you decided it, and what's actually been checked. Five failure modes shaped Edwin's design and its patent-pending approach, and we tested our answers to most of them in controlled experiments early in 2026.

1. Why isn't a confident answer a correct one?

A model's self-reported confidence can sit near the top on almost everything. In one test, Gemini rated 97.5% of 200 fact-checking claims at 0.95 confidence or above (Feb 2026). A score that sits that high on nearly everything tells you the model processed the claim thoroughly. It doesn't tell you the claim is true. A memory that stores whatever sounds confident will eventually hand your AI something wrong, stated with total conviction.

Edwin's answer: keep "how well was this analyzed?" and "is it true?" as separate questions. A second AI provider challenges stored facts in the background, and every memory carries an honest label: Survived challenge, Not confirmed, Disputed, or Not yet challenged.Evidence: checking facts and showing the AI those labels raised accuracy 5.8 points on STALE, a benchmark of questions about facts that changed. (pre-registered, Oct 2026)

The full explanation is in why AI confidence isn't truth.

2. When is a change not a contradiction?

When a chief executive leaves one company and joins another, two facts exist: "Jane Smith is CEO of Acme Corp" (then) and "Jane Smith is CEO of Newco" (now). That's an update, not a contradiction. A memory that doesn't track time either flags it as a conflict or overwrites the old value and loses the history. Budgets, owners and plans change all the time, so a useful memory has to tell an update from an error.

Edwin's answer: store every fact with the date it was true. When a budget goes from $500K to $750K, Edwin records the change as an update, not an error, and the old value stays on record with its date.Evidence: with source dates, Edwin classified 50 of 50 updates correctly; without them, 23 of 50. (internal test, Feb 2026)

3. How do you catch contradictions without crying wolf?

Some contradictions hide across several facts. "Alice manages Project X." "Alice is not in management." "Project X requires a management sponsor." Any two are consistent; all three together are impossible. A checker that compares facts only in pairs misses that. The opposite failure is just as costly: a checker that flags too much gets ignored, and then it protects no one.

Edwin's answer: read related facts together, so contradictions that span three or more facts can be spotted, and count a contradiction only when two different AI providers agree or you confirm it. One model's hunch isn't enough.Evidence: zero false alarms on a 150-pair test, catching 68 of 100 real contradictions. (internal test, Apr 2026)

4. What happens when the checkers start agreeing?

A second AI only helps if it can disagree with the first. Two copies of the same model tend to share the same blind spots, and a checker that always agrees adds cost without adding safety. Agreement means something only when disagreement was possible.

Edwin's answer: the challenger always comes from a different AI provider than the primary checker. Edwin monitors how far apart the models' judgments are, and the step that weighs their views is instructed not to average positions just to look balanced: when one model is confident and the other isn't, it can side with one of them.Evidence: published research by Zhang et al. found that mixing different models consistently improves debate among AIs. (arXiv:2502.08788, Feb 2025)

More on what works in using a second AI to check the first.

5. Who owns your AI's memory?

Memory built inside a single AI app lives on that company's servers, under its terms. If the product shuts down or the terms change, the history you built stays behind. As your AI learns more about your work, that history becomes more valuable, and more worth owning.

Edwin's answer: your memory is a SQLite file on your own Mac, with encrypted nightly backups to your iCloud Drive, and it works with any MCP-compatible AI client, so changing tools doesn't mean starting over.

Local-first, not local-only: to extract and check facts, Edwin sends text to AI providers (Venice, Google Gemini and Anthropic). Our privacy policy lists what goes where, and your AI memory should be yours covers the rest.

50 of 50
updates classified correctly with dates (23 without) · internal test, Feb 2026
0
false alarms, catching 68 of 100 contradictions · internal test, Apr 2026
+5.8 pts
accuracy from checking facts and showing labels · STALE, Oct 2026
~0.3%
of the tokens of pasting your whole history in · STALE, Aug 2026

What does this add up to?

A memory your AI can rely on: it knows when each fact was true, it says what's been checked, it counts a contradiction only when two providers agree or you confirm it, and it belongs to you. Edwin also stores decisions with the reasoning behind them, so you can ask why, not just what. And it's efficient: it answers from about 0.3% of the tokens of pasting your whole history in, at the same overall accuracy (STALE, Aug 2026, self-run).

We publish the full measurement record, including what didn't work, on our research page; the early experiments behind these five answers are in the experiment log.

Edwin is in private development.

Edwin remembers what you've told your AI: when it was true, why you decided it, and what's actually been checked. Send us a note and we'll tell you when it opens.