Kimetsu logoKimetsu
Memory Benchmark

Structured answerability (v2.8.0)

Paired evidence-delivery measurements, configuration and remaining gaps for the opt-in fact guard.

These are v2.8.0 opt-in evidence-delivery results, not a new overall BrainBench score or a claim about generated-answer accuracy. It adds source-bound configuration facts and reports which requested attributes are supported, missing or conflicting.

Released September 9, 2026: v2.8.0 is available. These measurements use implementation 3ae8329 and harness 2c74dad; they are not a new run of a tagged v2.8.0 binary or a v2.7.0-versus-v2.8.0 comparison.

v2.8.0 metric summary

Frozen structured-fact fixture: 45 cases, repeated twice.

MetricBaselineCandidateInterpretation
Unwanted injections15/18 (83.3%)3/18 (16.7%)80% fewer failures; 66.7 percentage points lower
Positive retrieval hits24/27 (88.9%)24/27 (88.9%)No positive-hit loss
Exact evidence metadataNot emitted36/45 (80%) per repeatNot generated-answer accuracy
Subsequent-query p95376.6 ms386.6 ms+10.0 ms (+2.7%)
Mean MCP response bytes613.5651.2+37.8 bytes (+6.2%)

Paired results — September 7, 2026

The campaign contains 688 observations over 299 distinct scenario/query cases. Both sides enable the fact guard; the baseline is the preceding answerability build and the candidate adds structured evidence. No positive-hit losses or execution errors were observed.

FixturePositive hits, old → newUnwanted injections, old → newSubsequent p95, old → new
Development: 210 queries169/197 → 169/1977/13 → 7/131,111.7 → 1,094.0 ms
Prior answerability: 44 queries24/24 → 24/240/20 → 0/20409.1 → 380.7 ms
New structured facts: 45 queries, two repeats24/27 → 24/2715/18 → 3/18376.6 → 386.6 ms

On the new fixture, 80% fewer unwanted injections means 15 failures became 3, or 83.3% → 16.7% of the 18 negative cases. Exact status, values, missing attributes and conflicts matched 36/45 cases in each repeat: 33/42 direct fact questions and all three broad controls. Rankings and metadata meaning were stable across repeats.

The new-fixture p95 increased about 2.7%, and mean MCP result size increased from 613.5 to 651.2 bytes (6.2%). Peak MCP working set remained about 653.5 MiB. These short runs do not establish zero overhead or a statistically significant speed change.

All runs used cached BGE-small embeddings, default thread settings and a 6,000 output budget. Development used TinyBERT with cutoff 0.30; the other fixtures used the optional quantized mMARCO reranker with cutoff 0.55. No model default was promoted. The byte budget includes serialization and is not billed tokens.

What changes for an agent

With Orchid staging gateway port is 7319. recorded, a recognized port-and-timeout question can return port evidence and identify timeout as missing. Production evidence cannot fill a staging request. Conflicting eligible values stay marked as conflicting even when a capsule cap or budget removes one source. Exact unit comparison treats 30 seconds and 30000 ms as equivalent.

On a build containing this change, enable it with:

kimetsu config set broker.explicit_fact_guard true

The default is false; set it back to false to disable it. Use v2.8.0 or a source build containing the merged changes; v2.7.0 packages lack this feature. An older warm daemon must be restarted with the updated build to use its new behavior.

Schema 15 maintains a rebuildable SQLite projection bound to claim revisions, source events, validity and exact evidence excerpts. Corrections and replay refresh the projection; the tagged agent record path is covered. Ordinary retrieval skips fact hydration when the guard is disabled, but writes still maintain the projection. No extra model calls are needed; evidence sent to an agent still consumes context tokens.

Remaining gaps and measurement limits

Six compound-question cases per repeat omitted port evidence: three retrieved nothing and three retrieved retries alone. Hit@4 counts the latter as hits, which shows why retrieval hit rate is not complete answerability. Three cases using the subject word Unknown fell outside the conservative grammar and returned unrelated evidence through legacy fallback. These account for the remaining unwanted injections.

The fixture is assistant-authored and repeats templates across three project names. It is not 45 independent task families, does not measure an LLM reader, and cannot certify corpus-wide conflict detection. No grammar, parameters or expectations changed after these results. Retrieval per requested attribute and safer unsupported-subject handling remain follow-up work.

The measured implementation passed 1,470 workspace tests (six ignored), 132 benchmark Rust tests, 18 Python tests and six release CLI/MCP probes. Separate v2.8.0 release-preparation validation passed 1,500 workspace tests, six ignored, plus 132 benchmark Rust tests after the dependency refresh. Release validation and security scope.

Source: full report, frozen fixtures, raw observations and hashes.

On this page