Loom

When two facts disagree, the newer one is usually right. Loom asks what happens when it is not.

PLATE VIILOOM

Loom, replayed

Engram 07 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

A loom at work. Facts are strung as warp; each vertical thread that carries a conflict has its winner’s weft pass over it and its loser’s under, and the shuttle crosses to lay the row. Height is support mass, and the cloth grows below the reed as the stream is written.
Experiment 7 of 10 · The fact with more support wins a conflict.
Is a fact that fits the rest of memory more likely to be right?

On FactConsolidation the newest fact wins by design, so recency is hard to beat there. Loom is the experiment that asks the harder question, and it needs a test that Tide cannot pass by counting.

Archived claim
6 of 6
hand-built cases; “useless at scale”
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 7 of 10 in the Engram series. ← Horizon · Engram overview · Coral →

Facts are not independent. “Person B lives in City R” sits among facts about where B works, where B’s family lives and what B did last year. If a new fact claims B lives somewhere else, most of that surrounding web disagrees with it, and the claim is weakly supported. If instead several other facts back the new claim, the old one is the odd one out. Loom measures support and lets the better-supported side win.

Loom · who has more support

Mechanism · drawn, not run
a conflict, and who has more supportA · “B lives in R”A′ · “B lives in S”conflict“B works in R”“B’s sister is in R”“B works in S”“B moved house in 2024”“B’s mail goes to S”support(A)support(A′)the better-supported side wins, not the later date
Two claims cannot both be right; each is supported by the facts that agree with it, and the better supported side wins, not the later date. An illustration, not a result.

Mutual-kNN means two facts count as agreeing only if each is among the other’s nearest neighbours, which limits how much support a fact that sits near everything can collect.

2. Where it came from

In April this was called protein binding: chunks attach to each other at shared-entity “sites”, and a contradiction disrupts the fold and earns low trust. Four scripts built it up. One stored small arithmetic webs (a division, a multiplication, an addition that agree) and asked whether an update that breaks the web loses. A second carved sites as it ingested facts. A third asked whether original facts sit in denser neighbourhoods than their updates. The project page’s one-line verdict on all of it: perfect in the lab, six for six, and useless at scale, because the locks were too vague.

The audit adds the reasons. The toy cases were small enough to tune. The author set the expected winner for each, and the gradual-evidence test counted a tie as a correct answer. Entities were capitalised words. The density test hard-coded example chunk indices. The one scored number for the family, “knit/loom”, was 42% against cosine’s 39%, in the one-feature test that returned more than k chunks.

3. Hypothesis

Stated before running, in two parts. First, on FactConsolidation, Loom at equal k does not beat Tide, because the dataset’s rule is that the later fact wins and support is a noisier proxy for that rule than the date is. I would be pleased to be wrong.

Second, on the gradual-override micro-test, Loom flips a belief only after a plausible amount of contrary evidence, and Tide never does: it flips on the first newer fact, whatever backs it. That is the property worth having in a memory that is sometimes told something false. This is the part of Loom that Tide cannot pass by counting.

4. Design

Support
The agreement mass of a fact: the summed similarity to the facts that are mutual nearest neighbours of it. Computed over the embeddings the Dense anchor uses.
Conflict
Two retrieved facts that are close in meaning and differ in content are a candidate conflict. The winner is the one with more support; the other is dropped from the result.
Equal k
Resolution removes losers and the list is refilled from the ranking to k. It never returns more than Dense.
The micro-test
A small fact web, a contradicting update, and 0 to 5 further facts that back the update. We record how many it takes to flip the answer. A tie is a miss, not a hit, and the winners are fixed by a rule written down before the run.
The ablation
Support term off: the same pipeline picking by date. It should land on Tide.

5. How it is tested

Loom is read against Tide as well as Dense, because on this dataset Tide is the right control. The micro-test is run separately from the benchmark: its output is a flip point per memory, with the number of cases and their construction rule published beside it. Hand-built cases are suggestive, not evidence, and the page will call them that.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Loom · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Loom (support by mutual-kNN agreement)0.9520.8850.30not runnot run
Loom, recency decides conflicts0.9520.8850.57not runnot run
Loom, no subject regex0.9520.8850.39not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
HippoRAG-v2 (published, MAB Table 3, GPT-4o-mini)n/an/an/a54.05.0
GPT-5-mini 400K, full context (published, MAB Table 3)n/an/an/a78.028.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisLoom’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If Loom does not beat Tide on FactConsolidation, that is the expected result and not a failure, provided the micro-test shows a graded response.
  • If the micro-test shows the same flip point whatever the support, support is not what drives the choice.
  • If Loom loses to the support-off ablation, the “binding” in the name has been a date all along.

Sources

  1. isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
    context · internal · engram/archive/baleen_old/benchmarks/isolated-ab.py · as of 2026-04
  2. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  3. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  4. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  5. context · public · as of 2026-06-28
  6. context · public · as of 2025-03-04
  7. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 7 of 10 in the Engram series. ← Horizon · Engram overview · Coral →

Loom is one of ten experiments in Engram. Its rows come from one run on the pilot sets.