Web

The fact you need often shares no words with the question. It shares a name with a fact that does.

PLATE IVWEB

Web, replayed

Engram 04 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

A spider’s web catching facts. The question snags its dense seeds like flies; the spider crosses the web along the threads of one hop through a shared entity, and the five facts it wraps are the ones returned.
Experiment 4 of 10 · One hop through shared entities finds the updated fact.
Can one hop through a shared name find the updated fact?

Multi-hop questions need a chain of facts, and only the first link looks like the question. A graph of shared entities is the usual answer, and the published systems that use one are among the best on this benchmark.

Archived claim
+3 pts
43% vs 39%, returned 2k not k
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 4 of 10 in the Engram series. ← Ring · Engram overview · Bloodhound →

Ask “what is the country of citizenship of the spouse of the author of Book A?” and the best match by wording is the fact about Book A. The fact that answers the question is about Person C and Country D, and shares no words with the question at all. Web follows the names: Book A leads to Person B, Person B to Person C, Person C to Country D.

Web · one hop

Mechanism · drawn, not run
querytop hit by cosine“The author of Book A is Person B.”Book APerson Bentity indexPerson B ▸ facts 12 · 301 · 388Book A ▸ facts 12 · 40one hopother facts thatshare an entitymerge with the hits,re-score, cut to kplaceholders: Book A and Person B stand for any pair of entities in a fact
A hit names its entities; an index of entities leads to every other fact that mentions one, and those are merged with the hits and re-scored.

One hop reaches the second link only. A three-link chain needs a second pass, and I will not add one until the single hop has been measured.

2. Where it came from

The archived entity walk took the cosine top-5, found the entities in them, hopped once through a graph of shared entities and added what it found. Entities were any capitalised words. It scored 43% against cosine’s 39%.

The audit found two problems. The function returned k extra chunks on top of the k it started with, so it handed the metric twice as many candidates as cosine had. And capitalised-word entities are crude: every sentence starts with one, and a surname on its own and the same person’s full name count as different entities. A third problem was shared with the rest of the round, the substring-in-any-chunk metric, which rewards size.

3. Hypothesis

Stated before running. With proper entities and the result cut to k, Web beats Dense on FactConsolidation multi-hop with a paired interval above zero, and does nothing on single-hop, where the question names the entity already.

The prior is cautious. Published multi-hop scores on this benchmark are very low for everything except a long-context reasoning model: the best retrieval system in the table reaches 7 on a hundred questions at the 262K context, and HippoRAG-v2, which is built on a graph of exactly this kind, reaches 5. Any gain here will be small against a floor that low, and I will describe it that way.

4. Design

Entities
Named entities or noun chunks from an off-the-shelf tagger, normalised so that a short and a long form of one name map together. They replace the capitalised-word rule.
The hop
From the top hits, collect every stored fact that shares an entity, one hop only. The union is re-scored against the query.
Equal k
The merged set is cut to k, so the hop can only change which k, never how many. The archived version’s extra chunks are gone.
Order
Ties follow the harness rule: toward the more recently written fact. A fact and its update share every entity, so the tie-break decides which one the hop returns first.
The ablation
The archived capitalised-word entities, at k. It tells us how much of any gain is the tagger and how much is the idea.

5. How it is tested

Tier 0 on LongMemEval_S, where the question types that need several sessions are the interesting ones (27 of the 100 pilot questions are multi-session), and tier 1 on both FactConsolidation sets. The comparison that matters is Web against Dense and against HippoRAG-v2’s published rows, with the caveat that those are at a different context length from the pilot.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Web · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Web (entity hop, tuned)0.8310.6530.64not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
HippoRAG-v2 (published, MAB Table 3, GPT-4o-mini)n/an/an/a54.05.0
GraphRAG (published, MAB Table 3, GPT-4o-mini)n/an/an/a14.02.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisWeb’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If Web at equal k does not beat Dense on multi-hop, the first round’s “+3” was the extra chunks.
  • If the tagger version beats the capitalised-word version but neither beats Dense, entity quality was a real flaw and still not the bottleneck.
  • If Web helps multi-hop and hurts single-hop by pulling in stale versions of the same fact, one hop needs the recency prior from Tide, and we will combine them only after both have been measured alone.

Sources

  1. isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
    context · internal · engram/archive/baleen_old/benchmarks/isolated-ab.py · as of 2026-04
  2. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  3. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  4. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  5. context · public · as of 2026-06-28
  6. context · public · as of 2025-03-04
  7. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 4 of 10 in the Engram series. ← Ring · Engram overview · Bloodhound →

Web is one of ten experiments in Engram. Its rows come from one run on the pilot sets.