Pheromone

The first round said randomness beat precision. It also let the ants return more answers than anyone else.

PLATE IIPHEROMONE

Pheromone, replayed

Engram 02 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

An ant field. Ants leave the nest, walk to the dense top hits and on along the similarity edges; afterwards their trails are laid in ink and begin to evaporate, so the strokes swell where reads agree and thin to dotted ghosts where they stop.
Experiment 2 of 10 · Reads leave trails that later reads follow.
Does a memory get better at finding things by being used?

This is the “reads change the memory” idea in its plainest form. The old test could not tell it from simply returning more candidates, so the question has not been asked yet.

Archived claim
+7 pts
47% vs 39%, but more than k returned
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 2 of 10 in the Engram series. ← Tide · Engram overview · Ring →

Think of the stored items as places and the similarity between them as paths. A query starts a few ants at the best matches. Each ant walks from item to item, preferring paths that are similar and paths that earlier ants marked. The walk ends, the useful paths get more marker, and all markers fade a little. Over many reads the memory grows well-worn routes to the things people keep asking for.

Pheromone · one read

Mechanism · drawn, not run
one read: ants hop on similarity × trailseedin the top kp(next=j | at i) ∝ sim^α · trail^βthe ant’s stepwalks thatend in kreinforce, then evaporatereads →trail strengthtrail ← (1 − ρ) · trail
An ant walks a graph of similarity edges, biased by the trails already laid; walks that end in the top k deposit more, and every trail then evaporates a little.

The trail is the part that is new. Without it this is a randomised neighbourhood search, which is a perfectly respectable thing to test, and the ablation below tests exactly that.

2. Where it came from

The archived ant was added to cosine top-5 in the first round’s one-feature-at-a-time test. It scored 47% against 39%, the largest gain of any feature, and the project page drew the conclusion that randomness outperforms precision for memory recall. A separate script, ant-dowser, carried the fuller version: ant colony optimisation with evaporation, softmax choice, and a “casting” step borrowed from tracking dogs.

The audit found three problems with that number, all in the code. The ant returned more than five chunks while cosine returned exactly five, so the metric rewarded the larger set; the win may be nothing more than recall at about ten against recall at five. The ant also warmed up on ten queries drawn from the very question set it was then scored on. And the metric was a substring test on the retrieved text, over about 180 questions, one seed, no interval. A three-to-eight point difference is inside that noise.

3. Hypothesis

Stated before running. At equal k, with trails built only from earlier reads in the stream and never from gold answers, Pheromone is not worse than Dense and improves on later queries within a memory that is read many times.

One consequence is already clear from the data. A LongMemEval question has its own haystack, so a trail has nothing to accumulate on: every question starts with an empty memory. The LongMemEval cells therefore measure the walk alone. The trail can only show up on MemoryAgentBench, where a hundred questions are read against one memory. I expect the walk to add little, and the question is whether the trail adds anything on top.

4. Design

The walk
Ants start at the best dense matches and hop along nearest-neighbour edges. The returned list is the best k of what the ants visited, ranked by similarity, so it can never be larger than k.
The trail
Marker on edges, raised on walks that ended among the returned items and decayed by evaporation. It is updated only from earlier reads in the same stream. No gold answer ever touches it.
Warm-up
If the trail is warmed up before scoring, the warm-up uses questions outside the scored set, drawn from a seeded split that the code asserts is disjoint from the pilot sets. The run record says which.
Seeds
The walk is random, so every number is the mean of five seeds with the range alongside.
The ablation
Trails off: the same walk with the marker held at zero. If this matches the full method, the trail did nothing.

5. How it is tested

The key comparison is Pheromone against its own trails-off ablation, because that isolates the idea. The second is Pheromone against Dense at equal k, with a paired test over the five seeds. We also plot the score of each successive query within one MemoryAgentBench memory, since “improves with use” is a claim about order and an average hides it.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Pheromone · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Pheromone (3-seed mean)0.9480.8820.40not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
HippoRAG-v2 (published, MAB Table 3, GPT-4o-mini)n/an/an/a54.05.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisPheromone’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If Pheromone with trails matches the trails-off walk, trails do nothing, and the first round’s “ants win” was a bigger candidate set.
  • If the walk alone beats Dense at equal k, there is something to learn about randomised neighbourhood search, but it is not memory, and the page will say so.
  • If scores do not rise over the query sequence on MemoryAgentBench, the reads-change-memory claim has no support from this mechanism.
  • If the trail helps only when warmed up on the scored questions, that is the confound we removed, and it is the answer.

Sources

  1. isolated-ab · the first round’s one-feature-at-a-time test: chrono bonus, entity walk, ring, ant, manifold and knit, each added to cosine top-5
    context · internal · engram/archive/baleen_old/benchmarks/isolated-ab.py · as of 2026-04
  2. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  3. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  4. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  5. context · public · as of 2026-06-28
  6. context · public · as of 2025-03-04
  7. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 2 of 10 in the Engram series. ← Tide · Engram overview · Ring →

Pheromone is one of ten experiments in Engram. Its rows come from one run on the pilot sets.