Ring
A filter that adds the same number to every candidate is not a filter. It only changes how many you keep.
Ring, replayed
The idea comes from how a short-lived memory is stabilised: recent items sit in a bounded zone where they reinforce, compete or fade before anything becomes durable. For retrieval, that means candidates must earn their place.
1. The idea
Experiment 3 of 10 in the Engram series. ← Pheromone · Engram overview · Web →
A normal retriever scores each item against the query once and takes the top k. Ring takes twice as many candidates and puts them in a ring. On each lap every candidate is re-scored, not only against the query but against the others in the ring, and the weakest fall out. What is left after the last lap is the answer. A candidate that matches the query but is isolated from everything else in the ring is the sort of item that should go.
Ring · laps
2. Where it came from
The archived ring took the top 2k by cosine, ran three laps that each added 0.02 to every candidate’s score, and kept k+2. It scored 44% against cosine’s 39% in the first round’s one-feature-at-a-time test. The theory note behind it describes the ring as the zone of plasticity, where recent excitations stay unstable until they reinforce, compete or dissipate.
The audit read the code. Adding the same constant to every candidate on every lap cannot change the order of anything, so the “survival” did no work. The only thing the ring changed was the size of the result: seven chunks against cosine’s five. A bigger set scores higher on a substring-in-any-chunk metric whatever is inside it. The gain was the extra two, one seed, about 180 questions, no interval.
3. Hypothesis
Stated before running. A re-scoring rule that depends on the other candidates, applied at exactly k, beats Dense on the multi-hop set, where the facts that answer a question support each other, and is neutral on the single-hop set, where one fact is enough.
The worry is the other direction: re-scoring against neighbours rewards redundancy, and a pile of near-duplicates will survive together. Conflict sets are full of near-duplicates, an old fact and its replacement, so the ring may keep both and drop the one fact that is different. That would be a clean negative result.
4. Design
- Candidates
- The top 2k by cosine, as archived. The ring starts at twice the size of the answer.
- The survival rule
- On each lap an item is re-scored against the query and against the mean of its neighbours in the ring, and the lowest scorers are dropped. Only the final k are returned. The mixing weight and the lap count are fitted on a seeded split that is disjoint from the pilot sets.
- Equal k
- The result is k items, never k+2. This is the fix that matters most, and the runner asserts it.
- Seeds
- Ties between equal scores are broken with a seeded choice, so the run uses five seeds and reports mean and range.
- The ablation
- The archived rule, a constant bump per lap, at k. It should reproduce Dense exactly, and if it does not the harness has a bug we need to find.
5. How it is tested
Tier 0 on LongMemEval_S and tier 1 on both FactConsolidation sets, with paired tests against Dense. The constant-bump ablation is a check on the harness as much as on the idea: any difference from Dense there is a measurement error, and it is the first number we look at.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Ring · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Ring (survival laps) | 0.940 | 0.897 | 0.42 | not run | not run |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| MemGPT (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
How to read thisRing’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If the constant-bump ablation differs from Dense at all, we fix the harness before reading anything else.
- If Ring at equal k does not beat Dense on the multi-hop set, the “+5 points” of the first round was the two extra chunks, and the ring has nothing to add.
- If Ring hurts on the single-hop set because it keeps old and new versions together, that is a finding about re-scoring conflict sets, and it goes in the write-up.
Sources
- isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- Archived theory note · memory as a dynamical field; reads and writes both perturb it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 3 of 10 in the Engram series. ← Pheromone · Engram overview · Web →
Ring is one of ten experiments in Engram. Its rows come from one run on the pilot sets.