Why isn't a model's confidence the same as truth?
When a language model rates its own answer, it is reporting how thoroughly it analyzed the claim. That's useful, but it's a different question from whether the claim is true. In one test, Gemini rated 97.5% of 200 fact-checking claims at 0.95 confidence or higher (Feb 2026). When nearly every claim scores near the top, the score can't separate the true ones from the false ones.
Search scores have the same blind spot. When a memory system retrieves a fact because it closely matches your question, the match says the fact is relevant. It says nothing about whether the fact is still correct.
How does Edwin keep the two questions apart?
Edwin treats confidence and truth as two separate measurements:
- How well was it analyzed? The models' own confidence, combined from more than one model.
- Is it true? An independent check by a different AI provider that tries to dispute the fact.
In Edwin's design the two are multiplied, so a weak result on either side pulls the whole score down. A claim that is well argued but probably false (0.85 × 0.20 = 0.17) ends up low; a claim that is well argued and checks out (0.85 × 0.90 ≈ 0.77) stays high. In an early test of this design on 50 hand-made facts, it raised effective accuracy from 56% to 88% (Feb 2026).
This separation is part of Edwin's provisional patent application, filed February 26, 2026; our patent summary covers the rest of the filing.
What does Edwin show beside each memory?
A number is easy to misread, so beside every memory Edwin shows a plain label that says what actually happened to it:
- Survived challenge: a second AI provider tried to dispute it and couldn't.
- Not confirmed: it was challenged and the support fell short.
- Disputed: the providers disagree about it.
- Not yet challenged: no second provider has checked it yet.
Edwin shows the label rather than hiding memories behind it, so you and your AI decide what to rely on. "Not yet challenged" doesn't mean wrong; it means untested, and the label says so plainly.
Do honest labels make answers better?
Yes, measurably. We tested it on STALE, a benchmark of questions about facts that changed over time, with the analysis plan written down before the run. Checking facts in the background and showing the answering AI those labels raised accuracy 5.8 points, and the labels carried most of that gain: removing them took away 4.1 points. A lighter check at question time, which looks at the facts an answer is about to use, raised accuracy 5.3 points on the same benchmark.
The full results are in the October 2026 measurements on our research page.
What should you ask of any AI memory?
Two questions cut through. Does it tell you what's been checked, separately from how confident it sounds? And can you see the evidence behind that? Edwin answers the first with a label on every memory, and the second with a public measurement record. For the other ways memory goes wrong, see five ways AI memory fails.