Engram
A memory should change when you read it as well as when you write it. Ten ways to try, and one test they all have to pass.
Most agent memory is written once and searched many times. Engram asks whether a read should leave a mark too, and whether any of that beats a plain keyword index when both return the same number of results.
1. Memory, not retrieval
Most memory for AI agents is a store with a search box. Write a record, embed it, and when a question arrives, return the nearest records. After a hundred questions the store is exactly what it was before them.
Engram began from a different premise, written down in April 2026 under the name Baleen: memory is a field, not a store, and what matters is how easy each thing is to reach. Two of the working axioms carry the whole programme. Reads and writes both change the memory: a write adds a disturbance and a read re-excites what it touches, and both alter what is easy to reach next time. And change is gradual: recent items sit in a bounded zone where they reinforce, compete or fade before anything becomes durable.
retrieval write ──▶ [ index ] ──▶ read
unchanged by either
memory write ──▶ [ memory ] ──▶ read
│ ▲ │ │
└─ marks ─┘ └─ marks ─┘
both leave a trace on what is easy to reach nextThat premise is a claim, and it has to be earned against the plainest alternative there is: a keyword index. So each of the ten ideas on this page is a mechanism, mostly borrowed from biology (ants, working memory, a coral reef, a bloodhound) or from geometry (a hyperbolic disc, a tree seven wide), plus one from language. Each is built to do exactly one thing and then measured.
The question is not whether a mechanism sounds like memory. It is whether it beats BM25 at the same k.
2. Ten mechanisms
The first five change how a read ranks what is already stored. The next four change how the facts are organised. The last changes what is written. Each has its own page with the same structure: the idea, where it came from and what was wrong with its first test, the hypothesis, the design, how it is tested, a results table, and what would change our mind.
Glyph is the one I would read first. It asks whether a memory can carry its own manual, so that a model which has never seen the notation can read it. It continues the line of work in MARCO?, where models from four families agreed on a coordinate once they were handed a shared page of landmarks. Glyph tries to take the page away.
3. Why we started again
The first round ran from April 2026. It produced a scorecard, and the scorecard produced claims: the ant kernel beat cosine by seven points, a coral reef scored 93%, a notation held memories in a third to a sixth of the tokens. I made those claims, and I am withdrawing them. Before writing any of it up I had the archive audited line by line against its own scripts and result files, with Claude doing the reading. The numbers did not hold. These are the reasons, and they are the reasons the new harness exists.
- Unequal k
- Cosine returned five chunks. The ant, the entity walk, the knit and the ring returned more: seven chunks for the ring, up to ten for the walks. The metric asked whether the answer appeared anywhere in the returned set, and a bigger set always scores higher. Some of the headline gains may be nothing but that.
- Warm-up on test questions
- The ant warmed up on ten queries taken from the question set it was then scored on. The bloodhound’s training scenarios overlapped its evaluation scenarios. A method that has seen the questions has not been tested on them.
- Substring scoring
- The standard metric was “a gold answer string appears in any retrieved chunk”. That is retrieval recall on the loosest possible rule, not the benchmark’s answer accuracy, and it was set beside published accuracies as if they were comparable.
- A test set the author wrote
- The 93% was fourteen questions of fifteen on thirty-five facts written alongside the system it flattered. The hand-built conflict cases for the loom were set up with their expected winners by the same person, and a tie counted as correct.
- No logs, one seed
- Almost every script printed its result and saved nothing. The numbers survive only in a hand-written scorecard, from a single seed, over about 180 questions, with no intervals. A difference of three to eight points sits inside that noise.
What the first round claimed, and what was behind it
| Claim | What the audit found |
|---|---|
| Ant kernel: 47% against cosine’s 39% | The ant returned more than five chunks to cosine’s five, and warmed up on ten of the questions it was then scored on. |
| Ring: 44% | The ring added one constant to every score, so no order ever changed. It kept seven chunks, not five. |
| Entity walk 43%, knit 42% | Each returned k extra chunks on top of k. Entities were any capitalised word. |
| Manifold: “−74% to +4%” | No log behind either figure. The one logged test, an untrained projection on synthetic clusters, scored 0.35 against cosine’s 0.88. |
| Coral: 93% | Fourteen of fifteen on a 35-fact set the author wrote. On MemoryAgentBench with an LLM answering it was 22% against 19%, from thirty questions per scenario. |
| Notation: 3 to 6 times fewer tokens | Never measured; the code counted whitespace-separated words. Cross-model reading rested on a few hand-picked questions. |
| Notation on MemoryAgentBench: 4 of 100 | The adapter bypassed the memory and judged answers with a home-made judge. It said nothing about the design. |
None of the first round’s numbers is carried forward. The archive stays as history.
Some of the first round’s habits were good and survive. Add one feature at a time and never stack them. Test whether quality climbs with use, not just whether it starts high. Publish the failures: the hyperbolic manifold was recorded as a failure at the time, and that was the right call. What changes is the harness, and the ideas are rebuilt to fit it.
4. The standard every experiment is held to
- Fixed, seeded pilot sets
- LongMemEval_S: 100 of the 500 questions, allocated in proportion to question type, seed 20261008, of which 83 are scored once the official skipping rules are applied (abstention questions and questions with no user-side evidence turn are not scored). MemoryAgentBench FactConsolidation, single-hop and multi-hop at 6K: all 100 questions of each. The sets are committed.
- Equal k
- Every method returns at most the same k ids, and the runner asserts it. Ties break toward the more recently written item, for every method.
- No training on test data
- Anything a method learns or warms up on comes from a seeded split disjoint from the pilot sets, saved with the run and asserted disjoint in code. Online adaptation inside one memory may use only earlier reads in the stream, never a gold answer.
- Official metrics
- Retrieval: recall_any@k and NDCG@k, as the LongMemEval code defines them. Answers: the official FactConsolidation prompt and the official substring-exact-match scorer, with replies cut to ten tokens to emulate the official output cap.
- Intervals and paired tests
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG) and paired tests against BM25 and dense retrieval on the same questions (paired bootstrap and McNemar). Stochastic methods run on five seeds, reported as a mean and a range.
- Published rows
- Other systems’ numbers are copied from MemoryAgentBench (arXiv 2507.05257) and LongMemEval (arXiv 2410.10813), with the table named in each row. We do not re-run them, and we do not average across contexts.
- Stated first, published either way
- The hypothesis is written before the run. A result that is no effect, or worse, is published as plainly as one that worked.
Two anchors sit under every comparison. BM25 is the keyword baseline, and our implementation reproduces the official LongMemEval flat-BM25 ranking exactly: the top-ten sets are identical on all 83 scored questions. Dense is bge-small, cosine, with the query prefix from the model card. If an idea cannot beat these at the same k, it has not shown anything.
Known deviationsThe official MemoryAgentBench run answers with GPT-4o-mini, at most ten tokens, temperature 0.7. We answer with Haiku through the command line, which can set neither, and Haiku writes its reasoning out, so the official scorer on the full reply is meaningless. We score the reply cut to its first ten tokens and label every such number, and any variant with a changed prompt is reported separately. LongMemEval’s retrieval numbers in the paper are for the larger M setting, or do not say which variant they are; S is much easier, so we compare only with the one published row whose variant is stated. The pilot MemoryAgentBench sets are 6K contexts, while the published numbers are at 262K.
5. Results
One row per experiment, our two anchors, and a short list of published systems. Each experiment’s own page carries a longer table: the method and its ablations, the anchors, and the published rows nearest its idea.
All ten experiments · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Tide (recency re-rank on dense, λ tuned on the training split) | 0.940 | 0.888 | 0.40 | not run | not run |
| Pheromone (3-seed mean) | 0.948 | 0.882 | 0.40 | not run | not run |
| Ring (survival laps) | 0.940 | 0.897 | 0.42 | not run | not run |
| Web (entity hop, tuned) | 0.831 | 0.653 | 0.64 | not run | not run |
| Bloodhound (trained kernel, 3-seed mean) | 0.201 | 0.138 | 0.27 | not run | not run |
| Horizon (trained Poincaré embedding, blended) | 0.952 | 0.888 | 0.35 | not run | not run |
| Loom (support by mutual-kNN agreement) | 0.952 | 0.885 | 0.30 | not run | not run |
| Coral (as shipped, seed 0) | 0.916 | 0.800 | 0.43 | not run | not run |
| Canopy (K=7 curated topics), n=16 | 1.000 | 0.895 | n/a | n/a | n/a |
| Glyph, read cold by a different model: 88.0 SubEM on FC-SH 6K, n=50 | n/a | n/a | n/a | n/a | n/a |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| HippoRAG-v2 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 54.0 | 5.0 |
| Mem0 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 18.0 | 2.0 |
| GPT-4o 128K, full context (published, MAB Table 3) | n/a | n/a | n/a | 60.0 | 5.0 |
PendingResults land when each run completes, and a page changes only when its row does. No cell of ours is filled: not the ten methods, and not our own BM25 and dense anchors. The published rows are real numbers from the two papers. They are at 262K context and the pilot is 6K, so 6K results will be reported beside this table, labelled, and not set against 262K rows. Canopy is a LongMemEval experiment only, so its MemoryAgentBench cells are n/a.
6. What this is not
- It is not a result. Until a cell is filled, nothing here says any of the ten works.
- The pilot sets are small: 83, 100 and 100 questions. Intervals will be wide, and differences of a few points will not be resolvable. Paired tests help, and they do not make a small set large.
- At 6K, single-hop FactConsolidation is close to the ceiling for a strong answering model, so it separates methods poorly. The multi-hop set is where most of the room is, and published scores there are very low for nearly every system.
- The published rows are other people’s pipelines, with other answer models and other judges. They are context, not a leaderboard.
- The ten were chosen from what the archive already held. They are not a survey of what memory could be, and the names are mine.
- It is not a product. Nothing here ships, and no number here is a benchmark claim until it has an interval.
Method and sources
ReproducibilityEvery experiment is one Memory plug-in with the same read and write interface, run through the same runner. Datasets are fetched at pinned revisions and checked by hash. The LongMemEval release used is the original one the paper’s tables came from, not the later cleaned release. Every answer-model response is cached under a hash of the model, system prompt and prompt, so a re-run is free and offline. Results are written per question, with ranked ids, so any number can be recomputed.
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- Archived theory note · memory as a dynamical field; reads and writes both perturb it
Engram is ten ideas about how a memory should behave, and one test they all have to pass. Nothing on these pages is a result yet, and they say so until it is.