Memory your AI agent can be held accountable for.
FerriteMem is a deterministic memory layer for AI systems, with no generative model on the write path or the read path. Nothing is summarised on the way in and nothing is generated on the way out. Every retrieval is reproducible, every result is cited to a stored record, and the whole engine runs inside your own boundary. Built for the places where “the AI said so” is not an acceptable answer.
The two empty slots are the whole difference. On the way in, a generative model that summarises or extracts “facts” can change what was said. On the way out, a model that rewrites the question can misread it, and one that phrases the answer can invent it. Removing both is what makes a result reproducible — and what makes it possible to show where the result came from.
What FerriteMem is
a memory that sits beside the model, and what that buys youMost “AI memory” lives inside a generative model — in weights and context windows you cannot inspect. FerriteMem takes the opposite position: a store of records and events that lives beside your model, which your model reads from and writes to under a human gate. The model reasons; FerriteMem remembers — verbatim, reproducibly, and with a citation for every result.
No generative model touches your records in either direction. Writing stores the record as it arrived — nothing is summarised, extracted or rewritten. Reading matches and ranks stored records with fixed search and ranking models that never generate text. What comes back is what went in, with its citation.
Reproducible measured
The same query returns the same answer, every time. No sampling, no drift, no model version changing underneath a result. Eight benchmark figures re-measured seven weeks later came back identical.
Auditable measured
Every result names the stored record it came from, the scope it was held under and the date attached to it. An auditor follows any answer back to its source instead of taking the system’s word for it.
Hard to poison by design
A hostile instruction hidden in a document is stored as text and never executed here: no generative model reads it on the way in or the way out, so neither the write nor the search can be steered by it. Your own model still reads what we return — that part stays yours. And nothing enters the lasting record without a person approving it.
Private by construction by design
Records never leave your boundary. No outbound call while answering, no telemetry, no update service that reads your data. Nothing is trained on what you store: the engine’s matching and ranking models are fixed and never updated from your data.
Works with your agent measured
FerriteMem supplies no model and does not care which you run — yours, your customer’s, or whichever you move to next year. It answers on demand and stays silent otherwise, so nothing is pushed into a context window on the chance it might be needed.
Token-free memory by design
The memory layer runs no generative model on write or read, so remembering and retrieving spend no generation tokens — and precise cited records mean your own model reads a tighter context too.
What is happening now
The live context of the moment — short-lived, and free to change as the work goes on.
What happened
Events and outcomes, kept word for word — the trail you can go back and check.
What was decided and known
The lasting record. It outlives every session and changes only through a person’s approval.
A fact you were taught, a thing that happened and a decision that was made are not the same. They live for different lengths of time, and not everyone should be able to change them the same way. FerriteMem keeps these kinds apart.
The metric that matters
finding something is easy; finding everything is the jobEvery memory system answers one of two questions, and they look almost identical. One asks whether anything relevant came back. The other asks whether everything did. On the same retrieval over the same records these give opposite verdicts — and the difference between them is the entire regulated market.
“Find me something”
the chatbot question
One relevant result is a win. Right for search and assistants: you asked, you got something useful. If a second relevant record existed and never surfaced, no harm done.
- one relevant piece found
- three others missed, unnoticed
“Find me everything”
the audit question
Every relevant result must surface. Miss one and the answer is wrong however many you found — because the missing piece is exactly the one a regulator, a court or a clinician asks about.
- three of four found
- the fourth decides the outcome
94.8% of the time, everything relevant comes back.
On a public long-memory benchmark across 500 questions, FerriteMem returns everything relevant 94.8% of the time, and something relevant 99.4% of the time. We lead with the strict figure — the one that fails when anything is missed — because it is the one that matters when an answer has to survive review.
The field optimises and publishes the left-hand panel, and that is the right metric for search and chat. Regulated work only pays for the right-hand one. Most systems do not report it at all.
Completeness is only half the problem, though. Retrieve all six records perfectly and the answer is still wrong if the system trusts the one that was later corrected — which is the next section.
Conflict resolution under development
results coming soonWhen two stored facts contradict, FerriteMem resolves which one is current by a fixed, stated rule — deterministically, with no generative model in the path, so the resolution is auditable, replayable, and with no prompt for an attacker to steer. A broader evaluation of this behaviour is in progress; the measured results will be published here when it completes.
Five guarantees, in one system
each is a property some memory layer offers — few offer more than oneThe memory layers now available for AI agents each tend to choose one property and build around it. One keeps a model out of the read path. Another isolates tenants. Another puts a person between the agent and the permanent record. Each is a defensible product on its own. What is uncommon is all of them in one system, each measured rather than asserted — and that is what regulated work requires, because a buyer who needs one of these usually needs the others.
| Guarantee | What it prevents | How it is held |
|---|---|---|
| No generative model on the write path or the read path | A hidden instruction steering the search or being rewritten into memory; a summary replacing what was said; token cost on every write and every recall | Nothing is extracted, summarised or generated on the way in. Retrieval is BM25, fixed dense embeddings, a cross-encoder and a stated rule |
| Records stored verbatim | A model’s paraphrase standing in for the source; an audit trail that cannot be checked against anything | What was stored is what is returned, with its identifier and its date |
| Isolation, not selectivity | A restricted record shifting the ranking, or appearing at a lower position, for someone who may not see it | Scope is enforced inside every query. A record outside the caller’s groups is never retrieved — measured at zero leaks under adversarial testing |
| A person between the agent and the lasting record | An agent quietly writing a permanent fact that nobody approved | The durable tier changes only by a person’s approval. Nothing an agent writes reaches it without that step |
| The same answer twice | A result that cannot be reproduced for the auditor who asks six months later | No sampling, no temperature, no model version underneath. Eight figures re-measured seven weeks apart, unchanged |
What this does not claim. A system that learns its retrieval strategy from outcomes can score higher on a benchmark than one that fixes it. We chose the fixed rule because a strategy that moves cannot be replayed, and replay is what the sectors we serve ask for. That is a trade, made deliberately, and stated.
Results
every number with its terrain named| What | Result | On | Status |
|---|---|---|---|
| Retrieval — recall_all / recall_any | 94.8% / 99.4% | public long-memory benchmark, full N=500, retrieval-scored | measured |
| Fast-path latency | 19 ms | deterministic recall; 210 ms on the reranked path | measured |
| Write throughput | 7.8 / s | 64 concurrent writers, zero errors; each write searchable before it returns | measured |
| Reproducibility | 8 of 8 | eight figures re-measured seven weeks later, same corpus and configuration, unchanged | measured |
| Cross-tenant leaks | 0% | adversarial isolation testing; scope enforced inside every query | measured |
| Independent evaluation | in progress | scored by the evaluator on held-out data, with their model and their judging | in progress |