Pheromone
The first round said randomness beat precision. It also let the ants return more answers than anyone else.
Pheromone, replayed
This is the “reads change the memory” idea in its plainest form. The old test could not tell it from simply returning more candidates, so the question has not been asked yet.
1. The idea
Experiment 2 of 10 in the Engram series. ← Tide · Engram overview · Ring →
Think of the stored items as places and the similarity between them as paths. A query starts a few ants at the best matches. Each ant walks from item to item, preferring paths that are similar and paths that earlier ants marked. The walk ends, the useful paths get more marker, and all markers fade a little. Over many reads the memory grows well-worn routes to the things people keep asking for.
Pheromone · one read
The trail is the part that is new. Without it this is a randomised neighbourhood search, which is a perfectly respectable thing to test, and the ablation below tests exactly that.
2. Where it came from
The archived ant was added to cosine top-5 in the first round’s one-feature-at-a-time test. It scored 47% against 39%, the largest gain of any feature, and the project page drew the conclusion that randomness outperforms precision for memory recall. A separate script, ant-dowser, carried the fuller version: ant colony optimisation with evaporation, softmax choice, and a “casting” step borrowed from tracking dogs.
The audit found three problems with that number, all in the code. The ant returned more than five chunks while cosine returned exactly five, so the metric rewarded the larger set; the win may be nothing more than recall at about ten against recall at five. The ant also warmed up on ten queries drawn from the very question set it was then scored on. And the metric was a substring test on the retrieved text, over about 180 questions, one seed, no interval. A three-to-eight point difference is inside that noise.
3. Hypothesis
Stated before running. At equal k, with trails built only from earlier reads in the stream and never from gold answers, Pheromone is not worse than Dense and improves on later queries within a memory that is read many times.
One consequence is already clear from the data. A LongMemEval question has its own haystack, so a trail has nothing to accumulate on: every question starts with an empty memory. The LongMemEval cells therefore measure the walk alone. The trail can only show up on MemoryAgentBench, where a hundred questions are read against one memory. I expect the walk to add little, and the question is whether the trail adds anything on top.
4. Design
- The walk
- Ants start at the best dense matches and hop along nearest-neighbour edges. The returned list is the best k of what the ants visited, ranked by similarity, so it can never be larger than k.
- The trail
- Marker on edges, raised on walks that ended among the returned items and decayed by evaporation. It is updated only from earlier reads in the same stream. No gold answer ever touches it.
- Warm-up
- If the trail is warmed up before scoring, the warm-up uses questions outside the scored set, drawn from a seeded split that the code asserts is disjoint from the pilot sets. The run record says which.
- Seeds
- The walk is random, so every number is the mean of five seeds with the range alongside.
- The ablation
- Trails off: the same walk with the marker held at zero. If this matches the full method, the trail did nothing.
5. How it is tested
The key comparison is Pheromone against its own trails-off ablation, because that isolates the idea. The second is Pheromone against Dense at equal k, with a paired test over the five seeds. We also plot the score of each successive query within one MemoryAgentBench memory, since “improves with use” is a claim about order and an average hides it.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Pheromone · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Pheromone (3-seed mean) | 0.948 | 0.882 | 0.40 | not run | not run |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| HippoRAG-v2 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 54.0 | 5.0 |
How to read thisPheromone’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If Pheromone with trails matches the trails-off walk, trails do nothing, and the first round’s “ants win” was a bigger candidate set.
- If the walk alone beats Dense at equal k, there is something to learn about randomised neighbourhood search, but it is not memory, and the page will say so.
- If scores do not rise over the query sequence on MemoryAgentBench, the reads-change-memory claim has no support from this mechanism.
- If the trail helps only when warmed up on the scored questions, that is the confound we removed, and it is the answer.
Sources
- isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 2 of 10 in the Engram series. ← Tide · Engram overview · Ring →
Pheromone is one of ten experiments in Engram. Its rows come from one run on the pilot sets.