All posts
Technical

Using a Second AI to Check the First: What Works

Published · Updated 4 min read
By Alexander Snyder, inventor of Edwin

A second AI is worth having when it is genuinely different from the first and its verdict reaches the answer as a clear label. What we learned building Edwin's cross-provider checks.

Multi-ModelVerificationBenchmarks
Short answer: A second AI makes memory more reliable when it comes from a different provider than the first, when real disagreement is allowed to count, and when its verdict reaches the answer as a clear label. That's how Edwin uses one: checking facts and showing the AI those labels raised accuracy 5.8 points on the STALE benchmark (pre-registered, Oct 2026).

Why use a second AI at all?

Every model makes mistakes it can't see. The appeal of a second opinion is the same as in any field: a reviewer who didn't write the draft catches what the author missed. Research supports the idea: Du et al. found that several model instances debating over rounds improved factuality and reasoning (arXiv:2305.14325 (opens in a new tab), May 2023). Building Edwin taught us that the details decide whether it works.

Lesson 1: the second AI has to be genuinely different

Two copies of the same model share the same training and often the same mistakes, so they tend to agree. Zhang et al., in "Stop Overvaluing Multi-Agent Debate" (arXiv:2502.08788 (opens in a new tab), February 2025), found that debate among copies of one model often fails to beat simpler methods, and that mixing different models consistently improves it.

In Edwin: the challenger always comes from a different AI provider than the primary checker, and it never falls back to the primary's provider.

Lesson 2: more rounds aren't more truth

It's tempting to assume that if one round of checking is good, three are better. Our own measurement says otherwise: on a 998-claim fact-checking benchmark, three verification rounds scored 4.3 points below one (pre-registered, Oct 2026). Research on AI teams points the same way: Pappu et al., "Multi-Agent Teams Hold Experts Back" (arXiv:2602.01011 (opens in a new tab), v4 May 2026), found that self-organizing teams of models underperformed their best member. So we judge a check by whether it improves answers, not by how much deliberation it adds.

Lesson 3: disagreement is the signal, so protect it

A checker that always agrees adds cost without adding safety. Edwin monitors how far apart the models' judgments are, and the step that weighs their views is instructed not to average positions just to look balanced: when one model is confident and the other isn't, it can side with one of them rather than split the difference.

In Edwin: a contradiction counts only when two different AI providers agree or you confirm it.Evidence: zero false alarms on a 150-pair test, catching 68 of 100 real contradictions. (internal test, Apr 2026)

Lesson 4: the value is in the label

The biggest gain didn't come from the challenge's own verdict. It came from telling the answering AI what had been checked. Every memory in Edwin carries one of four labels:

  • Survived challenge: a second AI provider tried to dispute it and couldn't.
  • Not confirmed: it was challenged and the support fell short.
  • Disputed: the providers disagree about it.
  • Not yet challenged: no second provider has checked it yet.

On STALE, a benchmark of questions about facts that changed, checking facts and showing the AI those labels raised accuracy 5.8 points. A lighter check at question time, which looks at the facts an answer is about to use, raised accuracy 5.3 points on the same benchmark.

+5.8 pts
accuracy from checking facts and showing labels · STALE, Oct 2026, 95% CI 2.4–9.1
+5.3 pts
accuracy from a check at question time · STALE, Oct 2026
2
different AI providers must agree before a contradiction counts
4
plain labels, one on every memory

How does Edwin's checking work today?

In the background, a primary model and a challenger from a different provider review stored facts a few at a time, and each memory's label updates as it is checked. The result is a memory that says plainly what has been tested and what hasn't. The full measurement record, including the experiments that shaped this design, is in the October 2026 measurements on our research page; the problems the checks are built to solve are in five ways AI memory fails.

Edwin is in private development.

Edwin remembers what you've told your AI: when it was true, why you decided it, and what's actually been checked. Send us a note and we'll tell you when it opens.