Bloodhound
To teach a search to find the answer, start at the answer and walk back to the question.
Bloodhound, replayed
If you know where the answer is, the path from the question to it is a free label. A policy of a hundred-odd parameters, trained on those labels, would be a nearly weightless way to search a memory it has never seen.
1. The idea
Experiment 5 of 10 in the Engram series. ← Web · Engram overview · Horizon →
Most search is told how similar things are and nothing about how to move. Bloodhound is the opposite: a walk through memory, one hop at a time, where a tiny network scores each possible next hop from a handful of features. Training is the unusual part. Start from a known answer, walk back toward the question, and call that path perfect. Every step on it is a positive example.
Bloodhound · training and use
The kernel never sees what a fact says, only how the step looks. That is what makes it portable, and what makes it a long shot.
2. Where it came from
The archive holds a family of scripts: a reverse bloodhound that trained a small network on perfect paths, an entity version with seven features, kernels of 29, 73, 145 and 273 parameters, a “pack” of four specialised dowsers feeding one ring, and a content-blind navigator of 145 parameters that decided whether to explore or backtrack. They ran on MemoryAgentBench conflict-resolution scenarios.
None of it left a result. The scripts print to the screen, and the scorecard lists the one-feature tests and none of these. The audit also found the training split was often the first four scenarios with evaluation on all eight, which includes the four it trained on; and the paths were built from gold answers, so any run that reused them for scoring had read the answer key. A network of a few hundred weights cannot learn much from thirty queries per scenario either. The timeline simply moves on, with no written conclusion.
3. Hypothesis
Stated before running. A kernel trained on paths from rows that are disjoint from the scored ones beats Dense at equal k on FactConsolidation, with a paired interval above zero.
My prior is against it. The features are mostly similarities that Dense already uses, the network is tiny and the training data is thin, so I expect a small gain or none. What I would like to learn is whether the gap between trained and untrained kernels is zero. That is a cleaner question than whether Bloodhound wins.
4. Design
- Paths
- On training rows only, the chunk that holds the gold answer string marks the end of a path, and the walk back to the question is the label. Gold answers are used here and nowhere else.
- The kernel
- A tiny multilayer perceptron over the seven features, about 145 weights, as the archive had it. No larger network, so that a gain cannot be put down to capacity.
- Training split
- MemoryAgentBench rows disjoint from the scored ones, seeded, saved as a file and asserted disjoint in code. The scored questions are never seen in training, and the kernel at test time sees no answers.
- Equal k
- The walk is bounded so that the returned list is k items, the same as every other method.
- Seeds
- Initialisation and walk order are random, so five seeds, mean and range.
- The ablation
- The same network untrained. If it scores like the trained one, the training added nothing.
5. How it is tested
Tier 0 on LongMemEval_S and tier 1 on both FactConsolidation sets. The comparison that carries the idea is trained against untrained, then trained against Dense. We also report the training and test losses, because a kernel that fits its training paths and no others is the first thing the audit would suspect.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Bloodhound · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Bloodhound (trained kernel, 3-seed mean) | 0.201 | 0.138 | 0.27 | not run | not run |
| Bloodhound, untrained kernel (control) | 0.012 | 0.013 | 0.11 | not run | not run |
| Bloodhound, walk re-ranked by cosine | 0.952 | 0.890 | 0.44 | not run | not run |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| Contriever (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 18.0 | 7.0 |
How to read thisBloodhound’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If trained and untrained kernels score alike, learning from reverse paths did nothing.
- If the trained kernel beats Dense on a training row and not on a held-out one, it memorised paths, and the old design was fitted to its own test.
- If it works on FactConsolidation and fails on LongMemEval, the kernel learned a feature of one dataset, not navigation.
Sources
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 5 of 10 in the Engram series. ← Web · Engram overview · Horizon →
Bloodhound is one of ten experiments in Engram. Its rows come from one run on the pilot sets.