Tide

In a list of facts where the newest one is the true one, the cheapest memory is a clock.

PLATE ITIDE

Tide, replayed

Engram 01 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

A shoreline washed by a tide. Every fact is a shell on the sand, the oldest at the left and the newest at the right. A read sends the tide across the fifty candidates, and the returned five lift off the sand and drift out on the water, as high as the recency prior they were given: a solid triangle was pulled up by the re-rank, a hollow one pushed down.
Experiment 1 of 10 · Later facts win.
How far does “newer wins” get you on its own?

FactConsolidation is built so that a later fact in the list is meant to override an earlier one about the same thing. A method that happens to prefer recent facts gets credit for understanding conflict. Tide prefers recent facts and does nothing else.

Archived claim
+3 pts
43% vs 39%, one seed, no interval
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 1 of 10 in the Engram series. · Engram overview · Pheromone →

Write a fact, then write a newer one that contradicts it. A plain retriever sees two near-identical sentences and ranks them by similarity, which makes the order between old and new close to a coin toss. Tide breaks the tie with the one thing the store already knows: the order things were written in.

Tide · the score

Mechanism · drawn, not run
write orderoldestnewestrecency(i) runs 0 → 1 along the axisscore(i) = cos(q, i) + λ · recency(i)one weight λ, fitted off the scored questionsthe same k as every other methoddense order → re-ranked1234512345solid: pulled up · dashed: pushed down
A recency axis feeds one added term; the re-rank pulls some of the dense top hits up and pushes others down, and the list is cut to the same k as every other method.

It is a control before it is a method. If Tide closes most of the gap between Dense and the best published system on FactConsolidation, then every later experiment that claims to resolve conflicts has to beat Tide, not Dense.

Every conflict-resolution claim in this series is measured against Tide.

2. Where it came from

The archived version was a chrono bonus: add 0.15 times a chunk’s position in the list to its cosine score. It sat in the first round’s one-feature-at-a-time test and scored 43% against cosine’s 39%, a gain of three points on about 180 questions, one seed and no interval. A second script went further and boosted chunks that shared entities with a later chunk. It left no result file.

What was wrong with the test is that the bonus favours the last chunk, which in this dataset is the right answer by construction. That is a prior about the benchmark, and the first round never reported it as one. It also never ran the control that would have shown how much of every other gain it was: cosine plus a recency prior was missing from the baselines, along with BM25 and a hybrid. The audit’s own guess is that it would have absorbed most of the three to eight points claimed for the other features. Tide is that missing baseline, built properly.

3. Hypothesis

Stated before running. Tide scores higher than Dense on FactConsolidation single-hop at equal k, with a paired 95% interval above zero. On the multi-hop set it helps much less, because there the missing step is finding the second fact, not choosing between two versions of one.

On LongMemEval I expect it to help the knowledge-update questions and be neutral or worse elsewhere, so the overall figure is a sum of opposite effects. The pilot has 15 knowledge-update and 27 temporal-reasoning questions out of 100, and we will report by question type for that reason.

4. Design

Score
Cosine from bge-small, the same dense scores as the Dense anchor, plus a weighted recency term.
Recency
Position in write order, scaled to run from 0 to 1. LongMemEval sessions are written in date order; FactConsolidation facts keep their serial position.
The weight
One number, fitted on a seeded split that is disjoint from the pilot sets, saved with the run and asserted disjoint in code. It is never tuned on a scored question.
Equal k
Dense’s top candidates are re-scored and the list is cut to k. Tide never returns more than Dense does. Ties go to the more recently written item, the same rule every method uses.
The ablation
The same re-rank on BM25 instead of dense. It separates the effect of recency from the retriever it rides on.

5. How it is tested

Tier 0 is retrieval: LongMemEval_S session recall@5, recall@10 and NDCG@10. Tier 1 is answering: Haiku reads the retrieved facts at k=10 and the official SubEM scores the reply cut to its first ten tokens, which emulates the official output cap. Tide is compared with Dense and BM25 on the same questions with paired tests, and with the published rows in the table below.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Tide · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Tide (recency re-rank on dense, λ tuned on the training split)0.9400.8880.40not runnot run
Tide, λ tuned on LongMemEval only0.9520.8850.39not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
GPT-4o 128K, full context (published, MAB Table 3)n/an/an/a60.05.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisTide’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If Tide beats Dense on FC-SH with a paired interval above zero and closes most of the gap to the best published retriever, the FC-SH score is mostly a recency score. We will say so on every later page in the series.
  • If Tide does nothing on FC-SH, either my reading of the dataset is wrong or the weight fit failed. We check the fit before believing either.
  • If it lifts knowledge-update questions and costs more elsewhere on LongMemEval, a recency prior needs a gate that applies it only when a question is about change. That is a different experiment, and we will name it as one.

Sources

  1. isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
    context · internal · engram/archive/baleen_old/benchmarks/isolated-ab.py · as of 2026-04
  2. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  3. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  4. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  5. context · public · as of 2026-06-28
  6. context · public · as of 2025-03-04
  7. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 1 of 10 in the Engram series. · Engram overview · Pheromone →

Tide is one of ten experiments in Engram. Its rows come from one run on the pilot sets.