FerriteMem
Active development. FerriteMem v2.0.0 is in testing and validation, and is entered in an independent public evaluation of agent-memory systems. Figures here are measured under the conditions stated beside them.

We build the memory, and we publish how we measured it.

FerriteMem is built by a small engineering practice working on one problem: giving AI systems a memory that can be checked. FerriteMem is the result — a deterministic memory layer with no generative model on the write path or the read path.

Why this exists

the problem that would not go away

The useful thing about an AI system is that it can read far more than a person can. The difficult thing is that when you ask it what it read, you get a fluent answer and no way to check it.

In most work that does not matter. In the work we care about it is the whole problem: a clinician needs the medication mentioned once and never repeated, a reviewer needs the document that waives privilege, an investigator needs the ownership change from three years ago. Not a summary of them.

So FerriteMem was built the other way round. The memory holds records verbatim and hands them back with a citation. It resolves contradictions by rule rather than by asking a generative model to choose, and it never lets a model summarise or rewrite a record on the way in. It runs inside your boundary and reaches nothing while it answers.

Everything follows from one decision: no generative model on the write path or the read path. That is a constraint, and it costs us things other systems can do. It also means the answer you get today is the answer you get next year, and that you can show anyone where it came from.

How we report

measured, in progress, by design — and nothing in between

measured

Measured on a stated benchmark, reproducibly, with the conditions written beside the figure. If we cannot say what it was measured on, it does not carry this tag.

in progress

Under way but not finished. Not built yet, or not yet scored — and we will not phrase it so that it sounds complete.

by design

It follows from the architecture. No generative model in the core means nothing generated to hallucinate — that is a consequence of the design, not a measurement of it.

The same three colours run through every figure on this site: teal where a number is a measured strength, ochre where a curve gives ground, and a deeper red at its floor. We publish limitations beside capabilities and the whole curve rather than its peak. If a figure ever appears without its conditions, that is a fault — please tell us.

The benchmarks we cite

public, third-party, and named — with what each one measures

Every figure on this site comes from a benchmark someone else built and published. We did not write the questions, we did not choose the answers, and we have no ability to influence what either contains. That is the point of using them, and it is why they are named here rather than described vaguely as “internal testing”.

LongMemEval — 500 curated questions embedded in user-assistant chat histories, testing five memory abilities: information extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention. Our recall figures come from the full 500-question set. The benchmark’s own authors report commercial assistants and long-context models losing around 30% accuracy as history grows. Wu, Wang, Yu, Zhang, Chang and Yu — LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory, ICLR 2025. arXiv:2410.10813. Publicly released, with code and data.

MemoryAgentBench — FactConsolidation — the conflict-resolution task, built from counterfactual edits and testing whether a system correctly prefers later information over earlier. A broader evaluation of conflict resolution on this task is in progress; the measured figures will be published here when it completes. Hu, Wang and McAuley — Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions, ICLR 2026. arXiv:2507.05257. Publicly released, with code and data.

What a self-run benchmark number is worth

Less than one produced by someone else. The questions are public, so a system can be tuned toward them; the scoring is run by us, so nobody has audited it. We report these figures with their conditions because that is the honest way to report them — not because they settle anything.

Which is why an independent evaluation is underway

Scored by the evaluator, on held-out data, with their answer model, their prompts and their judging. When that result exists it will appear here, and it will be worth more than everything on this page.

Where we are

stated plainly, including what is not yet proven

Independent evaluation in progress

FerriteMem v2.0.0 is entered in a public evaluation of agent-memory systems. Scoring is run by the evaluator on held-out data, with their answer model, their prompts and their judging — not ours. The result will be published here whatever it says.

Testing and validation

The engine is in active development. Every figure on this site was measured under the conditions stated beside it, on our hardware and on public benchmarks. Neither of those is your data, which is why the way to establish fit is a pilot.

Working with us