Coral
The headline was 93%. It was fourteen of fifteen questions on a thirty-five-fact set the author wrote.
Coral, replayed
Coral has several mechanisms at once, so a result for the whole says nothing about which part earned it. The experiment is the whole, then each part removed in turn.
1. The idea
Experiment 8 of 10 in the Engram series. ← Loom · Engram overview · Canopy →
Each fact is a polyp. It drifts toward the reef of related facts and settles there, with a probability that rises with affinity. Every cycle a living polyp lays down a ring, and a query feeds the polyps it retrieves. Polyps that go unfed for three lean cycles calcify: the surface dies, the structure stays. When a new fact contradicts an old one, the reef detects non-self. The old polyp calcifies, loses half its score, and keeps a pointer to its heir. A query that lands on the old polyp walks the pointer to the newest living version.
Coral · lifecycle
The thresholds are meant to come from the reef itself rather than from constants: a contradiction is declared when a rival’s similarity is an outlier against the reef’s own distribution, and a fact with many dependants resists being replaced.
2. Where it came from
Coral was the last design of the first round and the one with the most results, which is why it needs the most care. There were four numbers under one name. 93% was fourteen questions of fifteen on a thirty-five-fact synthetic scenario the author wrote, and the project page later labelled it an LLM evaluation of a MemoryAgentBench scenario. On MemoryAgentBench with text matching, Coral scored 20% to cosine’s 17% on one scenario. With an LLM answering, three scenarios gave 27, 17 and 23 against 23, 10 and 23 for cosine: a mean of 22 against 19, from thirty questions each, so one question was worth 3.3 points and the mean difference is about one question per scenario.
The audit adds to that. Coral retrieved thirty candidates and re-ordered them before keeping ten, while cosine kept a plain ten. The prompt told the model to prefer the most recent fact and presented the facts in order, which gives recency to any retriever for free. One scenario’s seven-point gain was credited to a faster index that let competition “fire over a wider sample”, so an engineering change moved the accuracy. Settling is random, with one seed. “No external constants” was overstated: the code has a maintenance cost, a 30/70 mass split, a settling exponent, two cut-offs and a minimum reef size. And nothing was ever switched off to see which part mattered.
3. Hypothesis
Stated before running. Coral as shipped, at equal k, beats Dense on FactConsolidation single-hop with a paired interval above zero, and does not beat Tide: most of what the reef does to a conflict is the date, reached by a longer road.
The ablations are where I expect to learn something. If heirs are the active ingredient, removing them should cost most of the gain. If calcification and the immune response matter only through the heirs they create, removing either alone should change little.
4. Design
- As shipped
- The reef code as it stands, with no re-tuning and its thresholds as they were. The only change is that a read returns exactly k, not thirty candidates cut to ten after a private re-order.
- No calcification
- Superseded polyps keep their full score.
- No heirs
- Calcified polyps stay in the ranking but no longer point to a successor.
- No immune response
- A contradicting fact never triggers the cascade that displaces the old one.
- Seeds
- Settling is random and the approximate index adds its own variation, so five seeds, mean and range.
- Order
- Facts are given to the reader in the same order for every method. The reef does not get a re-order the others do not.
5. How it is tested
Tier 0 on LongMemEval_S and tier 1 on both FactConsolidation sets, with paired tests against Dense and Tide. The published rows the reef is read against are Mem0, Cognee and MemGPT, the memory systems in the paper’s table, alongside BM25 and the plain embedding retriever. Their numbers are at a longer context than the pilot, so the comparison is indicative, and we will say so beside it.
The ablation table, one part off at a time, is the real product of this page. We will report it even when the whole does not beat Dense.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Coral · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Coral (as shipped, seed 0) | 0.916 | 0.800 | 0.43 | not run | not run |
| Coral, no calcification | 0.952 | 0.881 | 0.39 | not run | not run |
| Coral, no feeding | 0.916 | 0.800 | 0.43 | not run | not run |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| Mem0 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 18.0 | 2.0 |
| Cognee (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| MemGPT (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
How to read thisCoral’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If Coral as shipped does not beat Dense at equal k, the first round’s plus three points was a larger candidate set and a prompt that favoured recent facts.
- If it beats Dense and ties Tide, the reef is a recency prior with extra steps.
- If removing heirs costs more than removing calcification or the immune response, the supersession pointer is the idea worth keeping, and the biology around it is the story.
- If the five seeds disagree with each other by more than the gap to Dense, the reef is too random to say anything about.
Sources
- Coral reef memory architecture · polyps, reefs, calcification, immune response, heirs, self-calibrating thresholds
- The reef as built · the Polyp, Reef and Archipelago classes, its synthetic test and its LLM benchmark
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 8 of 10 in the Engram series. ← Loom · Engram overview · Canopy →
Coral is one of ten experiments in Engram. Its rows come from one run on the pilot sets.