Canopy

Let a model do the tidying overnight, and let the reader walk a tree no more than seven wide.

PLATE IXCANOPY

Canopy, replayed

Engram 09 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

A tree. The curator grows boughs for topics, seven wide, and hangs the facts as leaves. A read climbs from the trunk, the boughs it takes thicken, and the leaves it returns are ringed.
Experiment 9 of 10 · An offline curator grows a topic tree seven wide.
Is a tree built by a model easier to search than a flat list?

Hierarchies make long stores navigable for people. A curator model can build one at write time, once, and every later read pays less. Whether the reads also get more accurate is the open part.

Archived claim
1 of 20
vs 0 of 20, on different measures
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 9 of 10 in the Engram series. ← Coral · Engram overview · Glyph →

The store keeps the raw record as the source of truth and builds a tree on top of it. The bottom tier is the original entries. The middle tier groups them into topics. The top tier is a few short summaries. A curator model, running offline and slowly, does the grouping and writes the summaries. A reader starts at the top, picks which of at most seven children to open, and repeats until it reaches entries. Seven is borrowed from the hexagonal grid the idea is named after, where a cell has six neighbours.

Canopy · three tiers

Mechanism · drawn, not run
three tiers, seven wideL0summariesL1topicsL2raw entriescurate · offline, slowa model groups L2 into L1and writes L0 and L1at most 7 under eachnavigate · at query timedescend, then rank theentries reachedreturn the best kthe bold path is one descent
A slow offline curator groups the raw entries into topics and summaries, none wider than seven; at query time a read descends the tree and ranks the entries it reaches.

The tree is derived and disposable. Rebuild it and the grouping may come out differently, which the design accepts: a question about two topics should still find both by their shared ancestry whichever way the boundary falls.

2. Where it came from

The archived H3 was a curator (Haiku), three tiers of markdown nodes, and node ids made from a hash of the title so that a stable cluster keeps its identity across rebuilds. It was tested once, on the first twenty questions of one conversation from the LoCoMo benchmark, using only the first 200 turns. BM25 got none of them and H3 got one. Each query took about two minutes.

The audit found that the two numbers were not the same kind of thing. BM25’s score was whether the answer string appeared in the retrieved text, a recall measure. H3’s was whether a model’s written answer contained the gold, an accuracy measure. A BM25 score of zero for a keyword method on that data is odd enough that a matching bug is the likelier explanation. The match rule accepted a prediction contained in the gold answer, so a very short answer could pass. The “seven” was never enforced: the curator made five to fifteen clusters of any size. And the script’s docstring called its metric the benchmark’s official one, which it was not.

3. Hypothesis

Stated before running. Navigating a curated tree and then ranking the entries it reaches with the same dense scorer beats Dense alone at recall@5 on LongMemEval_S for the question types that ask about a topic across several sessions, and costs more per question.

The cost is not a footnote. Curating a haystack takes a model call per question, and a hundred questions means a hundred haystacks. I would count a method that needs minutes per query as unusable for an agent, and the table will say what each query cost.

4. Design

Curator
A language model builds the tree for each haystack once, with a fixed prompt and a hard cap on calls per run, written into the run record. The prompt is the archived one: summarise without inventing, name clusters by what they are.
Width
Seven, enforced this time. A node with more than seven children is split before the tree is used.
Navigation
The reader descends from the top and the entries it reaches are ranked by the dense scorer, so the output is a ranked list of at most k session ids.
Where it runs
LongMemEval_S, tier 0 only. The curation bill is why: it is paid once per haystack. The ~500-session LongMemEval_M run goes ahead only if the cost cap allows it. FactConsolidation is not run for Canopy.
Seeds
Curation is not deterministic, so if the budget allows it the tree is built twice per haystack and the two scores are reported.

5. How it is tested

The measurement is session recall@5, recall@10 and NDCG@10 against the official LongMemEval ground truth, which is a retrieval measure and the same one Dense and BM25 are scored on. The comparison is Canopy against Dense, paired, with the cost per question beside it. The published structure-augmented rows (RAPTOR, GraphRAG, MemoRAG) are in the table for context. They are MemoryAgentBench numbers, a different benchmark, which is why Canopy’s own MemoryAgentBench cells are n/a.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Canopy · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Canopy (K=7 curated topics), n=161.0000.895n/an/an/a
Dense on the same 16 questions0.9380.890n/an/an/a
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
RAPTOR (published, MAB Table 3, GPT-4o-mini)n/an/an/a14.01.0
GraphRAG (published, MAB Table 3, GPT-4o-mini)n/an/an/a14.02.0
MemoRAG (published, MAB Table 3, GPT-4o-mini)n/an/an/a21.07.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisCanopy’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If Canopy does not beat Dense at equal k, a tree built by a model has not made this store easier to search, and the cost buys nothing.
  • If it wins on recall and costs minutes a question, it is a result about cost, not usefulness, and the page will report it that way.
  • If two builds of the same haystack differ by more than Canopy differs from Dense, non-determinism is the finding.

Sources

  1. H3 hierarchical memory · three tiers, a curator model, K=7 neighbourhoods, non-deterministic by design
    context · internal · engram/archive/baleen_old/docs/architecture/h3.md · as of 2026-04
  2. The archived H3 benchmark · BM25 against curated navigation on one LoCoMo conversation, 20 questions
    measured · internal · engram/archive/baleen_old/benchmarks/results/h3_bench_0.json and h3_bench.py · as of 2026-04-26
  3. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  4. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  5. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  6. context · public · as of 2026-06-28
  7. context · public · as of 2025-03-04
  8. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 9 of 10 in the Engram series. ← Coral · Engram overview · Glyph →

Canopy is one of ten experiments in Engram. Its rows come from one run on the pilot sets.