Structured answerability (v2.8.0)
Paired evidence-delivery measurements, configuration and remaining gaps for the opt-in fact guard.
These are v2.8.0 opt-in evidence-delivery results, not a new overall BrainBench score or a claim about generated-answer accuracy. It adds source-bound configuration facts and reports which requested attributes are supported, missing or conflicting.
Released September 9, 2026: v2.8.0 is available.
These measurements use implementation 3ae8329
and harness 2c74dad; they are not a new run of a tagged v2.8.0 binary or a
v2.7.0-versus-v2.8.0 comparison.
v2.8.0 metric summary
Frozen structured-fact fixture: 45 cases, repeated twice.
| Metric | Baseline | Candidate | Interpretation |
|---|---|---|---|
| Unwanted injections | 15/18 (83.3%) | 3/18 (16.7%) | 80% fewer failures; 66.7 percentage points lower |
| Positive retrieval hits | 24/27 (88.9%) | 24/27 (88.9%) | No positive-hit loss |
| Exact evidence metadata | Not emitted | 36/45 (80%) per repeat | Not generated-answer accuracy |
| Subsequent-query p95 | 376.6 ms | 386.6 ms | +10.0 ms (+2.7%) |
| Mean MCP response bytes | 613.5 | 651.2 | +37.8 bytes (+6.2%) |
Paired results — September 7, 2026
The campaign contains 688 observations over 299 distinct scenario/query cases. Both sides enable the fact guard; the baseline is the preceding answerability build and the candidate adds structured evidence. No positive-hit losses or execution errors were observed.
| Fixture | Positive hits, old → new | Unwanted injections, old → new | Subsequent p95, old → new |
|---|---|---|---|
| Development: 210 queries | 169/197 → 169/197 | 7/13 → 7/13 | 1,111.7 → 1,094.0 ms |
| Prior answerability: 44 queries | 24/24 → 24/24 | 0/20 → 0/20 | 409.1 → 380.7 ms |
| New structured facts: 45 queries, two repeats | 24/27 → 24/27 | 15/18 → 3/18 | 376.6 → 386.6 ms |
On the new fixture, 80% fewer unwanted injections means 15 failures became 3, or 83.3% → 16.7% of the 18 negative cases. Exact status, values, missing attributes and conflicts matched 36/45 cases in each repeat: 33/42 direct fact questions and all three broad controls. Rankings and metadata meaning were stable across repeats.
The new-fixture p95 increased about 2.7%, and mean MCP result size increased from 613.5 to 651.2 bytes (6.2%). Peak MCP working set remained about 653.5 MiB. These short runs do not establish zero overhead or a statistically significant speed change.
All runs used cached BGE-small embeddings, default thread settings and a 6,000 output budget. Development used TinyBERT with cutoff 0.30; the other fixtures used the optional quantized mMARCO reranker with cutoff 0.55. No model default was promoted. The byte budget includes serialization and is not billed tokens.
What changes for an agent
With Orchid staging gateway port is 7319. recorded, a recognized port-and-timeout
question can return port evidence and identify timeout as missing. Production
evidence cannot fill a staging request. Conflicting eligible values stay marked
as conflicting even when a capsule cap or budget removes one source. Exact
unit comparison treats 30 seconds and 30000 ms as equivalent.
On a build containing this change, enable it with:
kimetsu config set broker.explicit_fact_guard trueThe default is false; set it back to false to disable it. Use v2.8.0 or
a source build containing the merged changes; v2.7.0 packages lack this feature. An older warm daemon must be restarted
with the updated build to use its new behavior.
Schema 15 maintains a rebuildable SQLite projection bound to claim revisions, source events, validity and exact evidence excerpts. Corrections and replay refresh the projection; the tagged agent record path is covered. Ordinary retrieval skips fact hydration when the guard is disabled, but writes still maintain the projection. No extra model calls are needed; evidence sent to an agent still consumes context tokens.
Remaining gaps and measurement limits
Six compound-question cases per repeat omitted port evidence: three retrieved
nothing and three retrieved retries alone. Hit@4 counts the latter as hits,
which shows why retrieval hit rate is not complete answerability. Three cases
using the subject word Unknown fell outside the conservative grammar and
returned unrelated evidence through legacy fallback. These account for the
remaining unwanted injections.
The fixture is assistant-authored and repeats templates across three project names. It is not 45 independent task families, does not measure an LLM reader, and cannot certify corpus-wide conflict detection. No grammar, parameters or expectations changed after these results. Retrieval per requested attribute and safer unsupported-subject handling remain follow-up work.
The measured implementation passed 1,470 workspace tests (six ignored), 132 benchmark Rust tests, 18 Python tests and six release CLI/MCP probes. Separate v2.8.0 release-preparation validation passed 1,500 workspace tests, six ignored, plus 132 benchmark Rust tests after the dependency refresh. Release validation and security scope.
Source: full report, frozen fixtures, raw observations and hashes.
BrainBench
BrainBench is Kimetsu's own reader-free capability benchmark: it drives the real binary across difficulty tiers and scores dedup, forgetting, importance, and calibration…
LongMemEval
How Kimetsu scores on LongMemEval, the public long-term-memory benchmark, with the exact setup so the number can be reproduced and compared.