FishMem

Measure the full memory loop.

FishMem publishes the harness, frozen configuration, artifact hashes, confidence intervals, and limitations so teams can challenge each claim and build their own release gates.

Full paired portfolio

The frozen 2026-08-24 run compares FishMem and mem0 OSS 3.1.2 on every paired item. FishMem scores 88.2% versus 83.8% on 500 LongMemEval oracle questions and 46.7% versus 41.0% on 400 BEAM 100K questions.

LoCoMo is the loss: across categories 1–5 and 1,986 paired questions, FishMem scores 67.5% versus mem0's 71.1%. That statistically significant negative result remains in the release decision.

  • LongMemEval: +4.4 pt, 95% CI +1.0 to +7.8
  • BEAM: +5.7 pt, 95% CI +1.8 to +9.5
  • LoCoMo: −3.7 pt, 95% CI −5.8 to −1.5
  • Same inputs and disclosed configuration within each pair

What these results do not prove

LongMemEval oracle is not LongMemEval-S and cannot support a human-beating claim. Recovered process attempts make total cost and end-to-end wall-clock comparisons non-publishable for this run.

FishMem also retrieved substantially more context than mem0 on LongMemEval and BEAM. The quality wins are real under the frozen protocol; a universal accuracy, speed, cost, or token-efficiency claim is not.

Use the harness, not the headline

Public datasets reveal categories of failure, but your users, write policy, scopes, providers, and answer criteria will differ. Freeze representative product traces, measure persisted state and downstream answers, and keep a held-out release set.

Inspect the benchmark runners