Bloodhound

To teach a search to find the answer, start at the answer and walk back to the question.

PLATE VBLOODHOUND

Bloodhound, replayed

Engram 05 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

A scent trail across contoured ground. The hound sets out from its flagged seeds and follows the walk hop by hop, leaving footprints darker where the learned kernel scored the step higher.
Experiment 5 of 10 · Learn to navigate by walking answers back to questions.
Can navigation be learned from where answers were found?

If you know where the answer is, the path from the question to it is a free label. A policy of a hundred-odd parameters, trained on those labels, would be a nearly weightless way to search a memory it has never seen.

Archived claim
No result
trained on scored scenarios; nothing saved
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 5 of 10 in the Engram series. ← Web · Engram overview · Horizon →

Most search is told how similar things are and nothing about how to move. Bloodhound is the opposite: a walk through memory, one hop at a time, where a tiny network scores each possible next hop from a handful of features. Training is the unusual part. Start from a known answer, walk back toward the question, and call that path perfect. Every step on it is a positive example.

Bloodhound · training and use

Mechanism · drawn, not run
training · answers knownanswerquerythe perfect path, read backwardsfeatures(step) → was it a good step?a tiny MLP, about 145 weightstest · no answers···querythe kernel scores each next hop from features alone · tick = score · keep kcosine · entity · temporal · trail · nexus · diversity · topology
Trained on paths walked back from known answers, a tiny kernel later scores each next hop from features alone, with no answers in sight.

The kernel never sees what a fact says, only how the step looks. That is what makes it portable, and what makes it a long shot.

2. Where it came from

The archive holds a family of scripts: a reverse bloodhound that trained a small network on perfect paths, an entity version with seven features, kernels of 29, 73, 145 and 273 parameters, a “pack” of four specialised dowsers feeding one ring, and a content-blind navigator of 145 parameters that decided whether to explore or backtrack. They ran on MemoryAgentBench conflict-resolution scenarios.

None of it left a result. The scripts print to the screen, and the scorecard lists the one-feature tests and none of these. The audit also found the training split was often the first four scenarios with evaluation on all eight, which includes the four it trained on; and the paths were built from gold answers, so any run that reused them for scoring had read the answer key. A network of a few hundred weights cannot learn much from thirty queries per scenario either. The timeline simply moves on, with no written conclusion.

3. Hypothesis

Stated before running. A kernel trained on paths from rows that are disjoint from the scored ones beats Dense at equal k on FactConsolidation, with a paired interval above zero.

My prior is against it. The features are mostly similarities that Dense already uses, the network is tiny and the training data is thin, so I expect a small gain or none. What I would like to learn is whether the gap between trained and untrained kernels is zero. That is a cleaner question than whether Bloodhound wins.

4. Design

Paths
On training rows only, the chunk that holds the gold answer string marks the end of a path, and the walk back to the question is the label. Gold answers are used here and nowhere else.
The kernel
A tiny multilayer perceptron over the seven features, about 145 weights, as the archive had it. No larger network, so that a gain cannot be put down to capacity.
Training split
MemoryAgentBench rows disjoint from the scored ones, seeded, saved as a file and asserted disjoint in code. The scored questions are never seen in training, and the kernel at test time sees no answers.
Equal k
The walk is bounded so that the returned list is k items, the same as every other method.
Seeds
Initialisation and walk order are random, so five seeds, mean and range.
The ablation
The same network untrained. If it scores like the trained one, the training added nothing.

5. How it is tested

Tier 0 on LongMemEval_S and tier 1 on both FactConsolidation sets. The comparison that carries the idea is trained against untrained, then trained against Dense. We also report the training and test losses, because a kernel that fits its training paths and no others is the first thing the audit would suspect.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Bloodhound · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Bloodhound (trained kernel, 3-seed mean)0.2010.1380.27not runnot run
Bloodhound, untrained kernel (control)0.0120.0130.11not runnot run
Bloodhound, walk re-ranked by cosine0.9520.8900.44not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
Contriever (published, MAB Table 3, GPT-4o-mini)n/an/an/a18.07.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisBloodhound’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If trained and untrained kernels score alike, learning from reverse paths did nothing.
  • If the trained kernel beats Dense on a training row and not on a held-out one, it memorised paths, and the old design was fitted to its own test.
  • If it works on FactConsolidation and fails on LongMemEval, the kernel learned a feature of one dataset, not navigation.

Sources

  1. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  2. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  3. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  4. context · public · as of 2026-06-28
  5. context · public · as of 2025-03-04
  6. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 5 of 10 in the Engram series. ← Web · Engram overview · Horizon →

Bloodhound is one of ten experiments in Engram. Its rows come from one run on the pilot sets.