Ring

A filter that adds the same number to every candidate is not a filter. It only changes how many you keep.

PLATE IIIRING

Ring, replayed

Engram 03 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

A running track. A read lines its pool up on the start; the weak drop out and lie down on the infield, crossed, while the survivors run the lap and take the podium in rank order.
Experiment 3 of 10 · Candidates must survive repeated re-scoring.
Can a candidate be tested before it is trusted?

The idea comes from how a short-lived memory is stabilised: recent items sit in a bounded zone where they reinforce, compete or fade before anything becomes durable. For retrieval, that means candidates must earn their place.

Archived claim
+5 pts
44% vs 39%, kept k+2 not k
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 3 of 10 in the Engram series. ← Pheromone · Engram overview · Web →

A normal retriever scores each item against the query once and takes the top k. Ring takes twice as many candidates and puts them in a ring. On each lap every candidate is re-scored, not only against the query but against the others in the ring, and the weakest fall out. What is left after the last lap is the answer. A candidate that matches the query but is isolated from everything else in the ring is the sort of item that should go.

Ring · laps

Mechanism · drawn, not run
a read, lap by laplap 0 · top 2k by cosinelap 1kfor each lapscore(i) = a · sim(q, i) + (1 − a) · mean sim (i, the others)the lowest fall outexactly k, never k + 2
The candidates are re-scored against the query and each other lap after lap; the lowest scorers fall out until exactly k remain.

2. Where it came from

The archived ring took the top 2k by cosine, ran three laps that each added 0.02 to every candidate’s score, and kept k+2. It scored 44% against cosine’s 39% in the first round’s one-feature-at-a-time test. The theory note behind it describes the ring as the zone of plasticity, where recent excitations stay unstable until they reinforce, compete or dissipate.

The audit read the code. Adding the same constant to every candidate on every lap cannot change the order of anything, so the “survival” did no work. The only thing the ring changed was the size of the result: seven chunks against cosine’s five. A bigger set scores higher on a substring-in-any-chunk metric whatever is inside it. The gain was the extra two, one seed, about 180 questions, no interval.

3. Hypothesis

Stated before running. A re-scoring rule that depends on the other candidates, applied at exactly k, beats Dense on the multi-hop set, where the facts that answer a question support each other, and is neutral on the single-hop set, where one fact is enough.

The worry is the other direction: re-scoring against neighbours rewards redundancy, and a pile of near-duplicates will survive together. Conflict sets are full of near-duplicates, an old fact and its replacement, so the ring may keep both and drop the one fact that is different. That would be a clean negative result.

4. Design

Candidates
The top 2k by cosine, as archived. The ring starts at twice the size of the answer.
The survival rule
On each lap an item is re-scored against the query and against the mean of its neighbours in the ring, and the lowest scorers are dropped. Only the final k are returned. The mixing weight and the lap count are fitted on a seeded split that is disjoint from the pilot sets.
Equal k
The result is k items, never k+2. This is the fix that matters most, and the runner asserts it.
Seeds
Ties between equal scores are broken with a seeded choice, so the run uses five seeds and reports mean and range.
The ablation
The archived rule, a constant bump per lap, at k. It should reproduce Dense exactly, and if it does not the harness has a bug we need to find.

5. How it is tested

Tier 0 on LongMemEval_S and tier 1 on both FactConsolidation sets, with paired tests against Dense. The constant-bump ablation is a check on the harness as much as on the idea: any difference from Dense there is a measurement error, and it is the first number we look at.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Ring · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Ring (survival laps)0.9400.8970.42not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
MemGPT (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisRing’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If the constant-bump ablation differs from Dense at all, we fix the harness before reading anything else.
  • If Ring at equal k does not beat Dense on the multi-hop set, the “+5 points” of the first round was the two extra chunks, and the ring has nothing to add.
  • If Ring hurts on the single-hop set because it keeps old and new versions together, that is a finding about re-scoring conflict sets, and it goes in the write-up.

Sources

  1. isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
    context · internal · engram/archive/baleen_old/benchmarks/isolated-ab.py · as of 2026-04
  2. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  3. Archived theory note · memory as a dynamical field; reads and writes both perturb it
    context · internal · engram/archive/baleen_old/docs/architecture/theory.md · as of 2026-04
  4. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  5. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  6. context · public · as of 2026-06-28
  7. context · public · as of 2025-03-04
  8. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 3 of 10 in the Engram series. ← Pheromone · Engram overview · Web →

Ring is one of ten experiments in Engram. Its rows come from one run on the pilot sets.