Why use a second AI at all?
Every model makes mistakes it can't see. The appeal of a second opinion is the same as in any field: a reviewer who didn't write the draft catches what the author missed. Research supports the idea: Du et al. found that several model instances debating over rounds improved factuality and reasoning (arXiv:2305.14325 (opens in a new tab), May 2023). Building Edwin taught us that the details decide whether it works.
Lesson 1: the second AI has to be genuinely different
Two copies of the same model share the same training and often the same mistakes, so they tend to agree. Zhang et al., in "Stop Overvaluing Multi-Agent Debate" (arXiv:2502.08788 (opens in a new tab), February 2025), found that debate among copies of one model often fails to beat simpler methods, and that mixing different models consistently improves it.
Lesson 2: more rounds aren't more truth
It's tempting to assume that if one round of checking is good, three are better. Our own measurement says otherwise: on a 998-claim fact-checking benchmark, three verification rounds scored 4.3 points below one (pre-registered, Oct 2026). Research on AI teams points the same way: Pappu et al., "Multi-Agent Teams Hold Experts Back" (arXiv:2602.01011 (opens in a new tab), v4 May 2026), found that self-organizing teams of models underperformed their best member. So we judge a check by whether it improves answers, not by how much deliberation it adds.
Lesson 3: disagreement is the signal, so protect it
A checker that always agrees adds cost without adding safety. Edwin monitors how far apart the models' judgments are, and the step that weighs their views is instructed not to average positions just to look balanced: when one model is confident and the other isn't, it can side with one of them rather than split the difference.
Lesson 4: the value is in the label
The biggest gain didn't come from the challenge's own verdict. It came from telling the answering AI what had been checked. Every memory in Edwin carries one of four labels:
- Survived challenge: a second AI provider tried to dispute it and couldn't.
- Not confirmed: it was challenged and the support fell short.
- Disputed: the providers disagree about it.
- Not yet challenged: no second provider has checked it yet.
On STALE, a benchmark of questions about facts that changed, checking facts and showing the AI those labels raised accuracy 5.8 points. A lighter check at question time, which looks at the facts an answer is about to use, raised accuracy 5.3 points on the same benchmark.
How does Edwin's checking work today?
In the background, a primary model and a challenger from a different provider review stored facts a few at a time, and each memory's label updates as it is checked. The result is a memory that says plainly what has been tested and what hasn't. The full measurement record, including the experiments that shaped this design, is in the October 2026 measurements on our research page; the problems the checks are built to solve are in five ways AI memory fails.