Web
The fact you need often shares no words with the question. It shares a name with a fact that does.
Web, replayed
Multi-hop questions need a chain of facts, and only the first link looks like the question. A graph of shared entities is the usual answer, and the published systems that use one are among the best on this benchmark.
1. The idea
Experiment 4 of 10 in the Engram series. ← Ring · Engram overview · Bloodhound →
Ask “what is the country of citizenship of the spouse of the author of Book A?” and the best match by wording is the fact about Book A. The fact that answers the question is about Person C and Country D, and shares no words with the question at all. Web follows the names: Book A leads to Person B, Person B to Person C, Person C to Country D.
Web · one hop
One hop reaches the second link only. A three-link chain needs a second pass, and I will not add one until the single hop has been measured.
2. Where it came from
The archived entity walk took the cosine top-5, found the entities in them, hopped once through a graph of shared entities and added what it found. Entities were any capitalised words. It scored 43% against cosine’s 39%.
The audit found two problems. The function returned k extra chunks on top of the k it started with, so it handed the metric twice as many candidates as cosine had. And capitalised-word entities are crude: every sentence starts with one, and a surname on its own and the same person’s full name count as different entities. A third problem was shared with the rest of the round, the substring-in-any-chunk metric, which rewards size.
3. Hypothesis
Stated before running. With proper entities and the result cut to k, Web beats Dense on FactConsolidation multi-hop with a paired interval above zero, and does nothing on single-hop, where the question names the entity already.
The prior is cautious. Published multi-hop scores on this benchmark are very low for everything except a long-context reasoning model: the best retrieval system in the table reaches 7 on a hundred questions at the 262K context, and HippoRAG-v2, which is built on a graph of exactly this kind, reaches 5. Any gain here will be small against a floor that low, and I will describe it that way.
4. Design
- Entities
- Named entities or noun chunks from an off-the-shelf tagger, normalised so that a short and a long form of one name map together. They replace the capitalised-word rule.
- The hop
- From the top hits, collect every stored fact that shares an entity, one hop only. The union is re-scored against the query.
- Equal k
- The merged set is cut to k, so the hop can only change which k, never how many. The archived version’s extra chunks are gone.
- Order
- Ties follow the harness rule: toward the more recently written fact. A fact and its update share every entity, so the tie-break decides which one the hop returns first.
- The ablation
- The archived capitalised-word entities, at k. It tells us how much of any gain is the tagger and how much is the idea.
5. How it is tested
Tier 0 on LongMemEval_S, where the question types that need several sessions are the interesting ones (27 of the 100 pilot questions are multi-session), and tier 1 on both FactConsolidation sets. The comparison that matters is Web against Dense and against HippoRAG-v2’s published rows, with the caveat that those are at a different context length from the pilot.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Web · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Web (entity hop, tuned) | 0.831 | 0.653 | 0.64 | not run | not run |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| HippoRAG-v2 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 54.0 | 5.0 |
| GraphRAG (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 14.0 | 2.0 |
How to read thisWeb’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If Web at equal k does not beat Dense on multi-hop, the first round’s “+3” was the extra chunks.
- If the tagger version beats the capitalised-word version but neither beats Dense, entity quality was a real flaw and still not the bottleneck.
- If Web helps multi-hop and hurts single-hop by pulling in stale versions of the same fact, one hop needs the recency prior from Tide, and we will combine them only after both have been measured alone.
Sources
- isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 4 of 10 in the Engram series. ← Ring · Engram overview · Bloodhound →
Web is one of ten experiments in Engram. Its rows come from one run on the pilot sets.