Kimetsu logoKimetsu
Memory Benchmark

How Kimetsu compares

How Kimetsu compares to mem0, Cognee, Zep, and Letta on the shared public benchmarks, and what we do not yet claim.

How Kimetsu compares to mem0, Cognee, Zep, and Letta on the shared public benchmarks, and what we do not yet claim.

How Kimetsu compares

mem0, Cognee, Zep, and Letta share a design: an LLM distills what to remember at write time, and most keep an LLM in the retrieval loop too. That buys accuracy at the cost of metered API spend on every question (mem0's own 2026 figures report ~7,000 tokens per retrieval call).

Kimetsu's memory pipeline makes zero LLM calls. Ingest, store, retrieve, and rerank are FTS5 + local embeddings + a local cross-encoder. The claim is not "more accurate"; it is the same accuracy band, without the LLM, the bill, or the cloud.

benchmarkKimetsu (local, model-free)mem0 (self-reported)Cognee (self-reported)
LongMemEval (_s)83.0% (200-q slice) · ~80.9% weighted94.4% (their reader + harness)not reported
LoCoMo (1,540 q)89.4%92.5%not reported
BEAM 100K73.3% (400 probes)n/a79%
BEAM 1M66.0% (300 probes)64.1% (700 probes)not reported
BEAM 10Mfuture work48.6%67%

Caveats, because the table is not apples-to-apples:

  • The 1M row uses the same token bucket, with different setups. Kimetsu's 66.0% covers 15 of 35 conversations; mem0 reports 64.1% over 700 probes. The readers and samples differ, so the score difference does not establish a head-to-head win.
  • Cognee leads at 100K/10M. Our 73.3% matches the prior public state of the art on 100K (the 0.735 Cognee cites as the number it beat), model-free. Cognee needs an LLM key on both the write and read paths.
  • Vendor numbers are self-reported and often do not reproduce (a published LoCoMo 91.6% re-ran closer to 58-66% in the 2026 roundups). We ship the exact harness and settings so ours can be checked.

Bottom line: the same accuracy band as the leading LLM-backed systems, with the entire memory pipeline local, free, and model-free. For a head-to-head, run your system through the same kbench harness.

Sources: mem0's 2026 benchmark roundup, Cognee's BEAM figures, the LongMemEval and BEAM papers.

Mem0's BEAM 1M figure was checked on September 12, 2026. Its current reported 64.1% replaces the 62% previously quoted here.

What we do not yet claim

  • Multi-hop retrieval of obliquely relevant memories is v2.6 work; the LongMemEval preference result (63%) is that ceiling on a public benchmark.
  • The LongMemEval number is a 200-question stratified slice with a specific reader, not the full 500.
  • BEAM covers the 100K and 1M buckets; 10M needs the write-time distiller in the loop and has not been run.
  • BrainBench's calibration track has the fewest scenarios and is still being scaled; read the scores per dimension.
  • Output-token savings in the ROI ledger are estimated, not metered.

On this page