Tide
In a list of facts where the newest one is the true one, the cheapest memory is a clock.
Tide, replayed
FactConsolidation is built so that a later fact in the list is meant to override an earlier one about the same thing. A method that happens to prefer recent facts gets credit for understanding conflict. Tide prefers recent facts and does nothing else.
1. The idea
Experiment 1 of 10 in the Engram series. · Engram overview · Pheromone →
Write a fact, then write a newer one that contradicts it. A plain retriever sees two near-identical sentences and ranks them by similarity, which makes the order between old and new close to a coin toss. Tide breaks the tie with the one thing the store already knows: the order things were written in.
Tide · the score
It is a control before it is a method. If Tide closes most of the gap between Dense and the best published system on FactConsolidation, then every later experiment that claims to resolve conflicts has to beat Tide, not Dense.
Every conflict-resolution claim in this series is measured against Tide.
2. Where it came from
The archived version was a chrono bonus: add 0.15 times a chunk’s position in the list to its cosine score. It sat in the first round’s one-feature-at-a-time test and scored 43% against cosine’s 39%, a gain of three points on about 180 questions, one seed and no interval. A second script went further and boosted chunks that shared entities with a later chunk. It left no result file.
What was wrong with the test is that the bonus favours the last chunk, which in this dataset is the right answer by construction. That is a prior about the benchmark, and the first round never reported it as one. It also never ran the control that would have shown how much of every other gain it was: cosine plus a recency prior was missing from the baselines, along with BM25 and a hybrid. The audit’s own guess is that it would have absorbed most of the three to eight points claimed for the other features. Tide is that missing baseline, built properly.
3. Hypothesis
Stated before running. Tide scores higher than Dense on FactConsolidation single-hop at equal k, with a paired 95% interval above zero. On the multi-hop set it helps much less, because there the missing step is finding the second fact, not choosing between two versions of one.
On LongMemEval I expect it to help the knowledge-update questions and be neutral or worse elsewhere, so the overall figure is a sum of opposite effects. The pilot has 15 knowledge-update and 27 temporal-reasoning questions out of 100, and we will report by question type for that reason.
4. Design
- Score
- Cosine from bge-small, the same dense scores as the Dense anchor, plus a weighted recency term.
- Recency
- Position in write order, scaled to run from 0 to 1. LongMemEval sessions are written in date order; FactConsolidation facts keep their serial position.
- The weight
- One number, fitted on a seeded split that is disjoint from the pilot sets, saved with the run and asserted disjoint in code. It is never tuned on a scored question.
- Equal k
- Dense’s top candidates are re-scored and the list is cut to k. Tide never returns more than Dense does. Ties go to the more recently written item, the same rule every method uses.
- The ablation
- The same re-rank on BM25 instead of dense. It separates the effect of recency from the retriever it rides on.
5. How it is tested
Tier 0 is retrieval: LongMemEval_S session recall@5, recall@10 and NDCG@10. Tier 1 is answering: Haiku reads the retrieved facts at k=10 and the official SubEM scores the reply cut to its first ten tokens, which emulates the official output cap. Tide is compared with Dense and BM25 on the same questions with paired tests, and with the published rows in the table below.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Tide · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Tide (recency re-rank on dense, λ tuned on the training split) | 0.940 | 0.888 | 0.40 | not run | not run |
| Tide, λ tuned on LongMemEval only | 0.952 | 0.885 | 0.39 | not run | not run |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| GPT-4o 128K, full context (published, MAB Table 3) | n/a | n/a | n/a | 60.0 | 5.0 |
How to read thisTide’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If Tide beats Dense on FC-SH with a paired interval above zero and closes most of the gap to the best published retriever, the FC-SH score is mostly a recency score. We will say so on every later page in the series.
- If Tide does nothing on FC-SH, either my reading of the dataset is wrong or the weight fit failed. We check the fit before believing either.
- If it lifts knowledge-update questions and costs more elsewhere on LongMemEval, a recency prior needs a gate that applies it only when a question is about change. That is a different experiment, and we will name it as one.
Sources
- isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 1 of 10 in the Engram series. · Engram overview · Pheromone →
Tide is one of ten experiments in Engram. Its rows come from one run on the pilot sets.