Evidence record · updated 2026-10-04

The evidence: how we measure Edwin

We publish the full measurement record — including what didn't work — because memory you rely on should be measured in public.

Every result below carries its date, its sample size and its replication status. The August 2026 STALE run and the October 2026 tests were pre-registered, and the harness, judge prompts and spend ledger are available on request.

Headline resultsOctober 2026 measurementsWhat didn't workHow we measure

Headline results

The strongest results in the record, each with its date, sample size and a link to the full detail below.

  • ~0.3%of the tokens, at the same overall accuracy

    Edwin answered from a mean of 541 input tokens against 160,535 for giving the model the whole history, with overall accuracy statistically indistinguishable (23.7% vs 22.8%). That is about 125× cheaper per question ($0.0012 vs $0.1491), and setup pays for itself after about 3 questions.

    STALE benchmark · Aug 2026 · 180 held-out scenarios · pre-registered, self-run

    The August 2026 cost result

  • +5.8points from checking facts and showing the AI the labels

    Checking facts and showing the AI those labels raised accuracy 5.8 points: 31.3 with the background checking cycle on against 25.6 with it off (95% interval +2.4 to +9.1). Removing the labels alone took away 4.1.

    STALE · Oct 2026 · 150 scenarios, 450 probes · pre-registered

    The verification and labels-stripped tests

  • +5.3points from one check at question time

    A single inexpensive check per question, run as the AI reads its memories, raised accuracy from 23.6 to 28.9 (95% interval +1.6 to +9.1), with no background cycle. The checks themselves cost $0.61.

    STALE · Oct 2026 · 450 probes · pre-registered · trade-offs in the table

    The read-time check and its trade-offs

  • 0false alarms in a 150-pair contradiction test

    Edwin's contradiction detector flagged nothing that wasn't a contradiction (precision 1.000) while catching 68 of 100 real ones (recall 0.680).

    Internal test (Exp 43) · Apr 2026 · 150 pairs

    Contradiction-detection results

  • 50/50updates classified correctly with source dates

    With the date each fact was true, Edwin told updates from errors in 50 of 50 cases; with the dates stripped, 23 of 50. Neither arm produced a false positive.

    Internal hand-written set (Exp 16–17) · Feb 2026 · 100 pairs

    The temporal-update experiments

  • 25.6 vs 13.1dated facts against plain search over the raw transcripts

    Turning meetings into dated facts beat plain keyword search over the raw transcripts, with the same answering model: −12.4 for search (95% interval −18.2 to −6.9), before any checking.

    STALE · Oct 2026 · 150 scenarios, 450 probes · pre-registered

    The keyword-search comparison

October 2026: pre-registered tests of today's pipeline

Run between 2026-09-30 and 2026-10-03, each under its own spend cap with its own ledger. Results are differences in percentage points with 95% intervals unless marked. STALE (meeting-style memory that goes out of date) and FEVER (Wikipedia fact-checking) are benchmarks, not anyone's real memory. Every October STALE row was answered by gemini-3.1-pro-preview; the August 2026 run below was answered by gemini-2.5-flash, so the two months' scores are not directly comparable. Every test is listed, whatever its outcome. Our write-up of the debate results: Using a Second AI to Check the First: What Works.

October 2026 measurements: date, test, sample size, result with interval, cost, outcome and reading, for every test run
DateTest · answering modelnResult [95% interval]CostOutcome
2026-10-01STALE: today's verification cycle on vs off, over the facts each answer readsAnswered by gemini-3.1-pro-preview

The gain reaches answers through the confidence labels shown beside the facts (see the labels-stripped test below). Every STALE fact is stale, so false doubts were not measured here.

150 scenarios, 450 probes31.3 vs 25.6: +5.8 [+2.4, +9.1]$40.66 (+$1.42 pilot)Helped
2026-10-01FEVER: three verification rounds vs oneModel: gemini-2.5-flash (first model)

Three rounds were less accurate. During the run the third model mostly fell back to the first model's own provider (Gemini), so this is not a clean test of mixing providers.

998 claims78.5% vs 82.8%: −4.3 [−6.2, −2.5]$4.00Hurt
2026-10-03FEVER: one bare gemini-2.5-flash callModel: gemini-2.5-flash

On world facts the pipeline's first step beat a bare call. The arms come from different runs (Oct 1 and Oct 3). In February a bare call had scored about 90%; with today's models it scored 74.5%.

998 claims74.5% [71.8–77.2]; vs one round −8.2 [−11.4, −5.0]; vs three rounds −3.9 [−7.4, −0.3]$0.94Helped
2026-10-01STALE on fresh scenarios: the current-state block on vs offAnswered by gemini-3.1-pro-preview

The block almost never said which value is current now: 9 of 2,008 attribute states. For scale: August's brains, answered by the same model on the 180 held-out scenarios, scored 31.7; the STALE paper's published table puts Gemini-3.1-pro reading the whole history at 55.2 (all 400).

150 scenarios, 450 probes25.8 [22.0–30.0] without the block, 23.8 with it: −2.0 [−4.9, +0.7]$87.43 (+$2.16 pilot)No measurable difference
2026-10-03STALE: one verification round vs threeAnswered by gemini-3.1-pro-preview

Three rounds beat one here, but only because one round produced no labels and no replacement candidates. That is a difference in what the labels say, not a measured difference in judgment — and the opposite of the FEVER result above.

150 scenarios, 450 probes25.6 vs 31.3: −5.8 [−9.1, −2.4]$5.27Mixed
2026-10-03STALE: the verification-on answers with the confidence labels stripped (replacement marks kept)Answered by gemini-3.1-pro-preview

The labels carry a measurable part of the gain. The replacement marks alone were not measurably better than nothing.

149 scenarios, 447 probeswith labels − stripped: +4.1 [+1.7, +6.6]; stripped − off: +1.7 [−1.1, +4.5]$10.46Helped
2026-10-03STALE: one cheap check per question at reading time, through the engine (no background cycle)Answered by gemini-3.1-pro-preview

It helped, at a cost: about one false “outdated” label in twenty on a fact that is current, and it caught only about a third of truly outdated facts. The pre-registered rule reads this as not settled. A first version of the check (same day) gave +4.4 [+1.3, +7.8] for $7.06.

150 scenarios, 450 probes; 1,406 facts for false doubts28.9 vs 23.6: +5.3 [+1.6, +9.1]; false “outdated” marks 5.3% [3.7, 7.1]$17.36 (the checks themselves $0.61)Mixed
2026-10-03STALE: plain keyword search (BM25) over the raw transcripts, same answering modelAnswered by gemini-3.1-pro-preview

Turning meetings into dated facts beat searching the transcripts, before any checking.

150 scenarios, 450 probes13.1 vs Edwin (verification off) 25.6: −12.4 [−18.2, −6.9]$7.47Helped
2026-10-03STALE: the answer window (the newer statement placed after the older one in what the model reads)Answered by gemini-3.1-pro-preview

The newer fact reached the answer more often (302 vs 234 probes), but the answers were not more often right. The window stays off.

150 scenarios × 3 replicatesoverall +0.2 [−3.6, +4.0]; state resolution −0.9 [−5.8, +4.0] (97.5% intervals)$40.92No measurable difference
2026-10-02Does agreement between two providers predict that a fact is wrong or replaced?

The second provider added no measured precision over the better single one. We keep it as a safety rule, not as evidence.

82 FEVER claims; 55 STALE replacements the benchmark can judgeshared doubt right 92.5% vs the better single provider 90.2%: +0.023 [0.000, +0.076]; agreed replacements correct 69.1% [55.1, 82.0]$0 (existing records)No measurable difference
2026-10-01Does the newer fact reach the answer? Edwin's top 6 facts vs keyword search over raw transcripts

Edwin reached it more often, with about a seventh of the text, but still missed the newer fact's conversation in 47% of probes.

180 scenarios, 540 probes53.1% vs 42.0%: keyword search −11.1 [−16.7, −5.6]$0Mixed

Read together: extraction carries the most weight; on STALE, checking helps through the labels an answer sees, not through anything stored; on FEVER, more rounds cost accuracy. What these tests do not yet cover is listed under how we measure.

August 2026: the cost result (STALE & BEAM)

We ran the STALE memory-invalidation benchmark (arXiv:2605.06527 (opens in a new tab)) on 180 pre-registered held-out scenarios with the paper's verbatim judge. Edwin and the control used the same answering model and the same judge, and the same prompt with one difference: only Edwin's carried the instruction to decline when unsure. The verification cycle was switched off in this run for cost; it was measured separately in October (above).

The headline is cost. Edwin answered from a mean 541 input tokens per answer against 160,535 for giving the same model the entire history — 0.337% — on corpora of about 150,000 tokens at August 2026 prices. From the per-call ledger: $0.4362 to ingest a corpus once, $0.0012 per query round, $0.1491 per round for full context (answer calls only; judging is left out on both sides). Break-even at 2.95 rounds. These token and cost figures are means over all 200 scenarios run; the accuracy figures are on the 180 held out. Overall accuracy was statistically indistinguishable: 23.7% against the control's 22.8%, a difference whose paired confidence interval spans zero. Full context did better on state resolution, the “what is true now” questions (detail under what didn't work).

On BEAM (arXiv:2510.27246 (opens in a new tab); 20 conversations of 100K tokens; exploratory, not pre-registered, no baseline in our harness), Edwin's overall score was 0.372. Its strongest category was abstention (0.65) — knowing when it doesn't know — and contradiction resolution scored 0.125. BEAM's temporal layer never actually ran: no entity carried a source date, and no “replaces” links formed in any of the 20 conversations, so this measured plain retrieval.

Splits on the held-out 180: Edwin scores 36.3% on same-attribute updates vs 11.1% on cross-attribute implicit invalidation. Purpose-built academic systems and raw long-context models score higher on overall recall. For context on the accuracy figure, 23.7% is a low absolute score: the STALE paper's published table puts Gemini-3.1-flash-lite, reading the whole history with no memory system, at 22.4, and the paper's own system at 68.0 (the published rows cover all 400 scenarios; ours, the 180 held out). The claim we make from this run is the cost, not the accuracy.

The papers behind Edwin's design

Five papers that shaped how Edwin is built. Each card gives the paper's finding, what we built from it and what we measured, with a status in words — including where our results differ from the paper's.

Key Finding

Multi-agent debate is often no better than simple single-model baselines under fair evaluation; using different models in the debate is the change the authors found to improve debate frameworks consistently.

Edwin Implementation

Edwin requires the challenger to come from a different provider than the first model (Gemini first; Venice GLM-5 as challenger, with Claude as its fallback). The third, reconciling model runs on the challenger's provider (Venice) and can fall back to Gemini.

Measured Result · Not reproduced

February 2026, 50 hand-made facts: mixed-model debate 88%, same-model debate 92%, single model 96%. October 2026, FEVER, pre-registered (n=998): three rounds 78.5% vs one round 82.8%, −4.3 points [−6.2, −2.5], with the third model mostly on the first model's provider during that run. Synthesis, which the extra rounds produced, measured 2% useful (1 of 56) and has been switched off since 2026-09-22.

Key Finding

Teams of LLM agents fail to match their expert member's performance, with losses of up to 41.1% on ML benchmarks (v4, May 2026), even when told who the expert is.

Edwin Implementation

Edwin's defer rule follows this idea: when one model's confidence is above 0.85 and the other's below 0.4, the reconciling model defers to the confident one instead of averaging. It is told not to split the difference.

Measured Result · Fires; effect untested

In the October FEVER run the reconciling model deferred on 230 of 453 decisions (144 to the first model, 86 to the challenger). No test yet separates what deferring changes from what it does not, so we cannot say whether it prevents the loss the paper describes.

Key Finding

MUSE uses Jensen-Shannon divergence to select well-calibrated subsets of LLMs for uncertainty estimation and combine their predictions. Disagreement between models carries information.

Edwin Implementation

Edwin uses the divergence for a different job: between its own first model's and challenger's confidence on a fact, as an online signal for which facts get re-checked first (independent claim 5 of the filing). MUSE uses it to choose which models to combine.

Measured Result · Negative result

Our earlier divergence was computed between the first and third models and was degenerate. Re-derived from the first model and the challenger (Sept 2026), it ranked wrong above right at AUROC 0.685 on 19 rows (exploratory); a pre-registered test in October found 0.41 and 0.56 on two sets, so disagreement does not mark a claim as false. Calibration of Edwin's confidence: AUROC 0.50 (n=51, Sept 2026), no better than chance.

Key Finding

83.6% on LongMemEval with an open-source 20B model (up from 39% for full context); 91.4% on LongMemEval and 89.61% on LoCoMo with a scaled-up backbone. The code (opens in a new tab) is MIT-licensed and self-hostable.

Edwin Implementation

Related Edwin test, graph-grounded vs isolated verification (Exp 14, Feb 2026, n=15): 10 of 15 contradictions vs 9 of 15 overall; 4 of 5 vs 3 of 5 on contradictions that need three or more facts.

Measured Result · Small samples

Our contradiction detector: an earlier F1 of 0.931 came from a self-written 200-pair set and is retired. Exp 43 (internal, 150 pairs, April 2026): precision 1.000, recall 0.680, F1 0.810. On third-party SNLI data (500 pairs, Feb 2026): precision 91.4%, recall 38.4%, F1 0.541. The live scan's precision is not yet measured.

Key Finding

Names five requirements for persistent agent memory: persistent storage, selective retention, associative routing, temporal chaining, and consolidation into higher-order abstractions.

Edwin Implementation

Requirement 1: a local SQLite store with encrypted nightly backups. Partial on 2, 3 and 4 (temporal chaining through episode tracking, 760 episodes of meetings and imports as of 2026-10-02). Requirement 5: synthesis is switched off (since 2026-09-22) and overnight consolidation records what it would change without acting (since 2026-10-03).

Measured Result · Our reading

Our own gap analysis, not the paper's: one requirement met, three partial, one switched off.

Three mechanisms built into Edwin

What each one does today, and what we have measured about it.

A Runtime Monitor on Model Agreement

A small control loop that notices when the models stop disagreeing.

Rolling 50-check window · automatic temperature bump · median divergence 0.073 (Sept 2026)

Multi-model review is only worth its cost while the models actually diverge. Edwin keeps a rolling window of the last 50 checks; when fewer than 15% end in disagreement it logs a warning, and three warnings in a row raise the challenger's temperature by 0.1, capped at 1.2, with a desktop notification. A separate check flags near-identical confidence between the models (a Jensen-Shannon divergence below 0.05) as an echo chamber.

It is a control loop, not self-awareness, and the 15% threshold is a judgment call rather than an empirically derived one. Measured between the first model and the challenger over 1,264 clean cross-provider cycles (Sept 2026), the median divergence is 0.073; an earlier reading compared the wrong pair of models (see retractions). The challenger and the reconciling model share a provider (Venice); since Sept 22 the reconciler's score no longer overwrites the challenger's. The nearest published work we found (arXiv:2605.24737 (opens in a new tab), May 2026) monitors LLM evaluators for compliance; Edwin's contribution is running such a monitor on a memory write path, not the idea itself (dependent claim 7c of our provisional filing).

A Local Store You Own

Local-first, not local-only: exactly what stays on the Mac and what leaves it.

Local SQLite store · AES-256 encrypted backups · testnet snapshots are a demo

The knowledge store is a SQLite file on your own machine. Copy it, back it up, open it with any SQLite client — no account, no API, no permission. Export runs over MCP to any compatible client. A nightly encrypted backup (AES-256, under your brain key) goes to your iCloud Drive.

What leaves the Mac: meeting transcripts go to Venice AI for fact extraction, in sections; facts are checked by Gemini and Venice, with Claude (Anthropic) as fallback and second-provider confirmer. Facts the extraction model marks sensitive get no further model calls and are never served to other AIs unless you opt in. The transcript they came from was already sent to Venice for extraction. The full list: what the Edwin software sends to other services.

Snapshots to Walrus are an opt-in demonstration, not a backup. Each one is encrypted, and its hash is recorded on a signed audit chain that the portal's Verify checks (since Sept 2026). The Ed25519 signature on the snapshot itself is computed but still not stored, and re-import does not verify it. Storage is the Walrus testnet, where copies expire in about three to four days (none are on mainnet): 49 of 50 recorded snapshots returned “not found” when we checked in September 2026. Verifiable agent memory is not a new idea: Merkle Automaton (arXiv:2506.13246 (opens in a new tab), 2025-06-16) and MemTrust (arXiv:2601.07004 (opens in a new tab), 2026-01-11) both predate our filing. More in Your AI Memory Should Be Yours.

A Log of Every Check

Every verification cycle leaves a record that can be audited.

14,306 outcomes to 2026-09-20 · 37.4% of the corpus rated useful

Every verification cycle records an outcome: which provider challenged a fact, what it said, and how the fact's confidence moved. Over time that becomes a record of which kinds of claims hold up and which trip the challenger.

The log held 14,306 outcomes to 2026-09-20. Most of them were recorded before Sept 22, when the challenger's verdict was still overwritten: in 96.24% of those cycles confidence went up, and in 94.6% the challenger applied a near-constant discount of about 0.2, so the early log says more about the pipeline's habits than about the facts. An internal quality audit puts 37.4% of the corpus at genuinely useful: extracted entities 50%, synthesized ones 2%, inferred ones 0%. Synthesis has been switched off since 2026-09-22. Whether the log becomes an asset is a measurement we have not made yet.

The labels on the live store, September 27, 2026: of 10,077 facts drawn from meetings, 25 had survived challenge, 71 were not confirmed, 46 were disputed and 9,935 were not yet challenged. Edwin shows the label beside each memory; it does not hide or gate memories by label.

Confabulation: measured in our own store, then switched off

In March 2026 synthesis was producing about 95 generated entities for every fact extracted from a source. A note dated 2026-03-03 named the failure mode, drawing on the reality-monitoring literature (Simons, Garrison & Johnson, 2017 (opens in a new tab)), measured the ratio, and specified caps on how deep and how much synthesis could go.

The caps cut the ratio to about 0.32 generated entities per extracted one by September 2026, counting synthesis and inference together (extraction 10,422 entities, synthesis 2,876, inference 489); synthesis alone was 0.276. Blind grading then found synthesis output 2% useful (1 of 56), so we switched synthesis off on 2026-09-22 and stopped serving generated facts on 2026-09-27.

Others have written about the same failure mode, among them SSGM (arXiv:2603.11768 (opens in a new tab), 2026-03-12) and “Honest Lying” (arXiv:2605.29463 (opens in a new tab), 2026-05-28), which calls it memory confabulation. We are not claiming to have been first.

Do labels on a memory change the answer? A June 2026 paper (“Manufactured Confidence”, arXiv:2606.29279 (opens in a new tab)) found that a passive “unverified” tag is ignored by the model that reads the memory. Our own October STALE test found the opposite for Edwin's labels: stripping them removed most of the verification cycle's gain (+4.1 [+1.7, +6.6]; the labels-stripped test in the October 2026 measurements). Different settings and different labels; both results stand.

The experiment log, February–April 2026

Selected results from the internal experiment log, none of them pre-registered, each with its status in words. The results that did not hold up are listed too, because they changed the architecture. The formula behind the veracity gate is explained in Why AI Confidence Isn't Truth, and What Edwin Shows Instead; the problems these experiments were aimed at are in Five Ways AI Memory Fails, and How Edwin Answers Each.

Replicated and single-run results

#01REPLICATED

FEVER Reproducibility

88.5% mean across 3 seeded runs (88.8 / 88.7 / 88.0, seeds 42/43/44, range 0.8 points), Feb 2026 — one of the four replicated results in this record.

What it measures: all three runs used one verification round — a single call to the first model, no challenge, no synthesis, with the block and cap thresholds tuned on seed 42. A separate February run on the same 1,000 claims compared rounds directly, on a different setup (Gemini first, a local 8B model as challenger, a meta-challenger; Feb 20–21): one round 87.7%, three rounds — the production setting then — 84.3%. In February a bare single API call (Exp 2) scored 89.8–90.3% over five runs on the same 1,000 claims; the filing discloses its 90.3%. Re-measured in October 2026 with today's models (n=998, pre-registered): one bare call 74.5%, one round 82.8%, three rounds 78.5%. So this number belongs to the first model, not to the verification pipeline, and we attribute it accordingly.

#09SINGLE RUN

Confidence Overconfidence Baseline

Gemini reported ≥0.95 confidence on 97.5% of claims and 0 of 200 landed in the [0.3, 0.7] band, mean 0.996. n=200, single model, Feb 2026.

This is the mechanism behind the ensembling result below: the confidence filter had no band to operate in. Measured on Gemini (not on 'GPT-4 class models', as our earlier writing said), and across all claims, not only wrong ones.

#13SINGLE RUN

Cross-Referential Contradiction Detection

Contradictions that only appear across three or more facts: pairwise comparison (Exp 12) found 0 of 5; the cluster check alone found 1 of 5; the integrated system (the cluster check plus later enrichment cycles) found 4 of 5. n=5, Feb 2026 — a direction, not a rate.

The filing attributes 1 of 5 directly to the cluster check, and the 4 of 5 to the integrated system. Since October 2026 the cluster check records what it would flag and changes nothing.

#15SINGLE RUN

Overnight Consolidation on Live Data

Two sweeps challenged 7 mature facts on live data; 6 were demoted and 1 survived. 2026-02-25.

There is no ground truth on whether the demotions were right: it shows the sweep runs on real memories, not that its calls are correct.

#16SINGLE RUN

Temporal Update Detection — internal 100-pair set

100% accuracy with source dates (100/100 claim pairs) vs 70% without — zero false positives across 50 temporal updates, 25 contradictions, 25 consistent pairs. Feb 2026.

The result that matters is the ablation: strip the dates and temporal-update classification falls from 50/50 to 23/50, with zero false positives in either arm. Context: the set is hand-written and kept in our own repository, the result has not been reproduced, and it is not the published DECODE benchmark (Nie et al., ACL 2021 (opens in a new tab)). 100/100 is likely a ceiling effect: on STALE, a third-party benchmark of the same problem, Edwin's overall score was 23.7% (Aug 2026). Time-aware, bi-temporal edge invalidation predates us in Zep's Graphiti (arXiv:2501.13956 (opens in a new tab), January 2025).

#18SINGLE RUN

Encrypted Snapshot Round Trip

A 132.5 KB snapshot was encrypted, signed, uploaded in 7.5 seconds, downloaded and re-imported, with every entity and edge identical afterwards. One run, Feb 2026.

A single round trip: it is a demonstration, not a durability or security test. Where the snapshot layer stands today is described under the local store above.

Unnumbered testSINGLE RUN

Graph Structure Against Random Null Models

Clustering coefficient 0.4843 on enriched-only entities: Z=61.64 against an Erdős–Rényi null, Z=19.14 against Barabási–Albert, p<0.001, N=162. Before the February 2026 filing.

Re-run after a reviewer argued the clustering was a bulk-import artifact; on enriched-only entities it went up, not down. It says the graph is not random. It does not say why.

#21SINGLE RUN

Adversarial Diversity Monitoring

A 15% disagreement floor was adopted as a health threshold: a judgment call from one run (2026-02-17), not a derived requirement, and never replicated.

The monitor is real and shipped. Between the first model and the challenger the median divergence is 0.073 (Sept 2026, 1,264 clean cross-provider cycles). Whether the floor is currently held is unmeasured; earlier readings from the wrong pair of models are withdrawn (see retractions).

#40SINGLE RUN

Few-Shot Extraction

F1 = 0.474 on 5 transcripts — deployed to production.

Production threshold, not academic ideal. Extraction is good enough to be useful, and in October it carried most of Edwin's lead over plain search; five transcripts is a small sample. An earlier internal comparison figure is retracted (see retractions).

Superseded, refuted or disproven

#03SUPERSEDED

Veracity Gate Value

56% → 88% effective accuracy on 50 hand-made facts (19 of 20 false facts caught, 5 of 30 true facts blocked), Feb 2026. Never replicated, and an internal review later called the effect expected: it compares no fact-checking against fact-checking.

Kept for the record, not as evidence. The live counterpart is the number that matters: as of 2026-09-20, 3 facts had been blocked across 14,306 production cycles, and in late September 2026 the gate ran on 0 of 1,071 cycles, because facts extracted from sources bypassed it by design. That lab-to-live gap is the largest we found in our own work. What changed: since 2026-10-01 observed facts go to the plausibility check (at most 40 a day), and since 2026-10-03 a doubt two providers share is recorded in shadow and changes no score.

#11NOT REPLICATED

Memory Retention over 30 Simulated Days

97% of 100 facts retained against 90% for a naive store, over 30 simulated days. Feb 2026.

A second run of the same configuration was a tie.

Unnumbered testLIKELY CIRCULAR

Counterfactual Challenge, Controlled Test

The challenge separated vulnerable from robust facts perfectly (Mann-Whitney U=225, p=1e-06). n=30, Feb 2026.

An internal review called this result likely circular or fragile, so we do not rely on it.

#17REFUTED AT SCALE

Temporal Detection Without Dates

Exp 17 is Exp 16 with the dates stripped: 70% overall (70/100), and 46% (23/50) on the temporal-update pairs. That 46% is the '23 of 50' quoted elsewhere on this site — it is Edwin's own pipeline without dates, not another system.

The claim that dated facts resolve this entirely is refuted at scale by our own STALE run: 1,632 'replaces' links were produced against roughly 17 real target pairs across 200 scenarios, and BEAM contradiction resolution scored 0.125. The temporal layer over-fires badly outside the 100-pair set, and that remains an open problem.

#19DISPROVEN

Ensemble Voting (Homogeneous Debate)

Ensemble voting across 3 models scored 92.5% vs 93.5% for the first model alone — a point lower, n=200, no confidence interval. Run 2 (Exp 20): 0 overrides triggered across 200 claims. Feb 2026.

No help, possibly a small cost. It led us to require a challenger from a different provider. That change has not been shown to help either: our February debate test found mixed-model debate no better than same-model (88% vs 92%, n=50), and in October three rounds again lost to one on FEVER (−4.3 points).

#42CIRCULAR

Zone Classification Accuracy

100% zone accuracy reported (150 synthetic facts, 500 live), Apr 2026.

Circular: it scored the classifier against labels produced by the same rule, so it shows the code follows its rule, not that the zones are right, and it measured the older three-zone classifier. Whether zone labels are right on real memories is unmeasured.

#44NO EFFECT

Evidence Grounding

Evidence grounding changed accuracy by −0.01, at three times the latency. n=100, Apr 2026.

No gain on this test, at a real cost in speed.

A selection. We don't quote a total: our own records state it five different ways (34, 35, 39, 41, 44), and we would rather say that than pick the flattering one. The provisional filing cites 22 experiments from a program it numbers at 35, two of them (19 and 20) disclosed as disproven. The full log, including the negative results and the two artifacts that were overwritten on 2026-09-20, is available on request.

What didn't work

The results that went against us, in one place. Each links to its full detail, and each changed what we built or how we describe it.

  • State resolution, August 2026

    On the “what is true now” questions, full context beat Edwin: 42.2 against Edwin's 25.0 (p = 9.7×10−5). The August run

  • More verification rounds on world facts

    On FEVER, three verification rounds scored below one. October table

  • Two October changes with no measurable effect

    The current-state block and the answer window; both stay off. October table

  • Agreement between two providers

    It added no measured precision over the better single provider; we keep it as a safety rule, not as evidence. October table

  • Disagreement as a signal of error

    Model disagreement does not mark a claim as false, and Edwin's confidence score has not yet been shown to be calibrated. The MUSE card

  • Synthesis

    Generated facts measured 2% useful; synthesis is switched off. Confabulation

  • Dated facts outside the test set

    On STALE the temporal layer over-fired badly. Exp 17

  • Evidence grounding and zone accuracy

    Web evidence in the veracity check changed accuracy by −0.01 at three times the latency (Exp 44); the reported 100% zone accuracy was circular (Exp 42).

  • Retention and consolidation tests

    A retention gain tied on its second run (Exp 11), and the controlled consolidation test was likely circular (the controlled test).

  • Ensemble voting, and the veracity gate in production

    Voting did not beat the first model alone (Exp 19); the gate's lab result did not carry over to live use (Exp 3).

Retractions and superseded results

Figures we published or relied on and later withdrew, and why. The line-by-line audit of our August figures is in the reproducibility tables; superseded experiments are marked in the experiment log.

The 13× false-premise figure: retracted

In the August STALE run Edwin rejected false premises far more often than the control (22.2% vs 1.7%). We have since classified every one of those passes, and roughly four-fifths of the gap is a refusal artifact: the pass criterion is satisfied by declining to answer, whether or not the system detected the staleness, and the control never received the refusal instruction. On probes where Edwin answers substantively the effect is about 5× (8.6% vs 1.7%). We previously published this as 13×; that framing was wrong, and we retract it here rather than quietly dropping it.

The August Mem0-OSS comparison: void

The Mem0-OSS comparison in the original August report is void and should not be cited, including by us. Its ingestion silently dropped a mean 26.5 of 50 sessions per scenario — 53.0%, in 59 of 59 scenarios, from 1,563 extraction-parse failures — so it answered from under half its input, and the counter meant to check that its retrieval was fair stopped at 20 because the call used a default limit of 20 results. That the starved baseline still reproduced its published score is itself evidence the harness was not measuring what it claimed.

Model-divergence readings before September 2026: withdrawn

Until September 2026 the divergence monitor compared the first model with the third (median 0.004) instead of with the challenger. Readings built on that pair (mean divergence 0.0118, 99.4% of entities below the coherent bar), the “echo chamber” reading and the conclusion that raising the challenger's temperature did not help are withdrawn. On 2026-09-20, 36.9% of checks had run with no challenger at all. Re-derived between the first model and the challenger over 1,264 clean cross-provider cycles, the median is 0.073 (Sept 2026).

Defects in the live system, found and fixed

On September 20, 2026 we found 437 entities whose value was the string “null” scoring 0.8 or higher; they were fixed the next day. Until September 27, 2026 our dashboard reported a headline share that counted any stored score as a pass, which overstated what had been checked; it now shows the label on each memory. Until September 28, 2026 a repair function rebuilt the snapshot audit chain on boot, which would have hidden an alteration; since then the chain's signatures are checked and never rewritten, and a break is listed for the owner instead.

“+26.8% F1 from few-shot extraction”: retracted

A ‘+26.8% F1 from few-shot’ figure circulated internally in which 0.374 is named both the winning score and the baseline thirteen lines apart, and the comparison run's output is missing. That figure is retracted; 0.474 on 5 transcripts (Exp 40) is the figure we have, and its comparison run is missing.

Reproducibility: published scores and re-runs

Each row is a published memory-benchmark number and a later re-run or correction, with who produced it, its source and the date we checked it. The published and re-run figures are not always for the same version. This is a record of reproducibility, not a performance comparison: we have not run another memory product under our harness (our one attempt is void). A longer, dated survey: How to Read an AI Memory Benchmark: A Buyer's Guide.

Published memory-benchmark scores and later re-runs or corrections, with who produced each, sources and the date each row was checked
System & benchmarkPublishedRe-run or correctedChecked
Mem0 · LongMemEval93.4, Mem0's earlier figure (now 94.4 · mem0.ai (opens in a new tab), April 2026 algorithm)73.8, a re-run of the 93.4 · maximem.ai (opens in a new tab), a rival vendor, on its own harness; published 2026-06-05 (updated 2026-09-26)2026-10-03
Mem0 · LoCoMo92.5 · mem0.ai (opens in a new tab), April 2026 algorithm81.08 for Mem0's open-source v2 (RAG 78.31, full context 77.16 in the same run) · arXiv:2601.07978 (opens in a new tab), v3 2026-05-30 (v5 2026-09-03), independent2026-10-03
Zep · LoCoMo~84 · Zep75.14 ± 0.17 (10 runs) · Zep's own corrected figure, 2025-05-12, after its ~84 was re-scored to 58.44 by Mem0's CTO (a competitor, 2025-05-08), in getzep/zep-papers issue #5 (opens in a new tab)2026-10-03
Graphiti, Cognee · LoCoMo—55–56 · arXiv:2601.07978 (opens in a new tab), v3 2026-05-30, independent2026-10-03
Edwin · STALE held-out (Aug 2026)23.7 · self-runnot re-run by anyone else; harness kept in our private repository, available on request—

Our own August figures under audit

An internal audit on 2026-09-20 recomputed 25 figures from our August report from the raw records. 19 reproduced exactly (two of them, the 13× premise ratio and the Mem0 score, reproduce as numbers but are retracted or void above). Six did not: three were wrong, three were stale text.

The six figures from our August 2026 report that did not reproduce exactly, and what the audit found
Figure as we stated itWhat the audit foundStatus
“Mem0 stored 20 memories in every scenario”20 was the default limit of the call used to count; the counter could not see the dropped sessions.Wrong
Clean subset, n=53: Edwin 20.1 vs Mem0 8.2, “confidence intervals disjoint”Disjoint only if every probe is independent. Resampled by scenario, the intervals overlap ([11.3–29.6] vs [2.5–14.5]).Wrong
BEAM: “temporal anchoring ran on content only”Materially understated: no entity carried a source date and no replacement links formed, in 20 of 20 conversations.Wrong
Edwin split by scenario type: 36.0 / 12.3These were the figures for all 200 scenarios; on the held-out 180 they are 36.3 / 11.1.Stale text
Control split by scenario type: 27.3 / 17.7Same mix-up; on the held-out 180 they are 27.4 / 18.1.Stale text
Methods note: “headline metric = held-out 380”The held-out set is 180 scenarios (the run was cut for budget).Stale text

How we measure

Pre-registration
The August 2026 STALE run and the October 2026 tests were pre-registered, each October test under its own spend cap with its own ledger. The February–April 2026 experiments are internal and were not pre-registered; the experiment log says so on every card.
Replication
Four results in this record have replicated. Most of the rest ran once. We give the replication status of each result, because in this field that is the number that matters.
Self-run benchmarks
Benchmark results are self-run; no official submission process exists for STALE or BEAM.
What you can check
The harness, judge prompts, pre-registrations, raw transcripts, per-call ledgers and each result file with its pre-registration commit are kept in our private repository and are available on request.
Not measured yet
False doubts in the background cycle, whether any of this holds on a real person's memory, calibration, and accuracy over time (a test against Mem0 is designed and pre-registered but not run). No experiment has measured whether divergence-driven scheduling saves cost or improves accuracy, and whether zone labels are right on real memories is unmeasured.
  • 4results replicated
  • 2disproven, published
  • 19/25August figures reproduced exactly under audit
  • 6that did not (3 wrong, 3 stale text)
  • $122.49ledgered spend, Aug 2026

Check our work

Edwin is in private development. The full experiment log — raw data, methodology and replication status for every result — is available to anyone who intends to check it.

Patent pending. Provisional patent application filed February 26, 2026: nine independent claims (among them effective confidence, graph-grounded verification, temporal update detection, data sovereignty and sleep consolidation) and fifteen dependent claims. Non-provisional due February 26, 2027. See our summary of the patent filing. Questions about what is implemented today: ask us.