Horizon

The first attempt bent the space with a function nobody trained, and it lost by a wide margin. That was the wrong test of the right question.

PLATE VIHORIZON

Horizon, replayed

Engram 06 · live replay · MAB FactConsolidation, 455 facts
Loading the replay

The Poincaré disc, tiled as in Escher’s Circle Limit prints: every pentagon is the same size to someone living there. Play the learn frames to watch training spread the facts out from the origin; whether general memories end up central and niche ones at the rim is the experiment. A query is inserted and joined to its neighbours by geodesics.
Experiment 6 of 10 · General memories sit at the centre, niche ones at the edge.
Does the shape of memory space matter?

Hierarchies grow exponentially and flat space is bad at holding them. Hyperbolic space is built for the job. The question is whether a learned hyperbolic space helps recall for conversation and fact memories, or only for textbook trees.

Archived claim
0.35 vs 0.88
P@5, hyperbolic vs cosine, synthetic
Now scored on
83 · 100 · 100
LongMemEval-S, FC-SH, FC-MH pilot questions
Held to
Equal k
no training on scored data
Result
Pending
filled in from one file

1. The idea

Experiment 6 of 10 in the Engram series. ← Bloodhound · Engram overview · Loom →

In flat space, the number of points that fit at a given distance grows polynomially. In a hyperbolic disc it grows exponentially, which is how a tree behaves. Put a tree in the disc with its root at the middle and the leaves near the rim and distances come out right: siblings are close to each other, cousins are far. A memory is not exactly a tree, but “general at the centre, niche at the edge” is a hierarchy.

Horizon · the disc

Mechanism · drawn, not run
the Poincaré discgeneralnichedistancestretchestoward the rim,room for manyniche memoriesfar from each otherbut all near thegeneral one theybelong to
In the Poincaré disc distance stretches towards the rim, leaving room for many niche memories that sit far from each other but near the general one they belong to.

The embedding is trained: it is fitted so that memories that were similar stay close and memories that were general end up near the origin. Retrieval ranks by hyperbolic distance.

2. Where it came from

The archived manifold took ordinary embeddings and squashed them into the ball with a tanh function. Nothing was learned. On synthetic clusters, retrieval at five scored 0.348 for hyperbolic distance and 0.296 with “dynamics” added, against 0.876 for cosine. Over twenty cycles of use the hyperbolic version stayed near 0.21 and cosine near 0.95. The project’s own note says it plainly: a naive projection hurts retrieval badly, and cosine is the right baseline until there is a hyperbolic embedding model.

So the honest result was negative, and also uninformative about hyperbolic space itself: an untrained projection is a strawman. The later variants (a PCA reduction to eight dimensions, an encoder trained by write-read-verify) left no result files, and the project page that advertised “−74% to +4%” has no log behind either number. The synthetic data was also built so that cluster membership was the ground truth, which makes cosine close to an oracle.

3. Hypothesis

Stated before running. A Poincaré embedding trained on the similarity graph of each memory does not beat Dense at equal k on these benchmarks, but norm (distance from the origin) tracks how general a memory is, and a small generality prior on top of hyperbolic distance helps on LongMemEval questions that ask for the gist of a topic.

The first half is a prediction that Horizon will not win, made so that the result cannot be spun. If it does win, we will have been wrong in a useful way.

4. Design

Embedding
A Poincaré embedding in the style of Nickel and Kiela, fitted by Riemannian gradient descent to the similarity graph of one memory’s items (nearest neighbours under bge-small). It is trained per memory, from the memory’s own contents and no labels.
Retrieval
Rank by hyperbolic distance from the embedded query. A norm-based generality term is added only in the second analysis, and its weight is fitted on a split disjoint from the pilot sets.
Equal k
The same k as every method.
Seeds
Training is stochastic: five seeds, mean and range.
The ablation
The archived tanh projection, run on the same real questions at last. It shows what the untrained version scores on the data that matters.

5. How it is tested

Tier 0 on LongMemEval_S and tier 1 on FactConsolidation, against Dense at equal k. The one figure specific to Horizon is a plot of embedding norm against a measured generality (how many other items an item is a near neighbour of), which shows whether the geometry puts general things in the middle at all, separate from whether it helps retrieval.

The standing rules for every Engram experiment apply here unchanged.

  • Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
  • Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
  • No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
  • Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
  • 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
  • Stochastic methods run on five seeds; the table reports the mean and the range.

6. Results

Horizon · results

Pilot sets · equal k · single run unless marked
MethodLongMemEval-S recall@5LongMemEval-S NDCG@10MAB FC-MH 6K answer in top 10MAB FC-SH 262K SubEMMAB FC-MH 262K SubEM
Horizon (trained Poincaré embedding, blended)0.9520.8880.35not runnot run
Horizon, hyperbolic distance only0.9280.7830.14not runnot run
Horizon, Euclidean control0.9520.8770.41not runnot run
BM25 (ours)0.9280.8750.2775.00.0
Dense bge-small (ours)0.9520.8850.3972.00.0
BM25 (published, MAB Table 3, GPT-4o-mini)n/an/an/a48.03.0
Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini)n/an/an/a28.03.0
RAPTOR (published, MAB Table 3, GPT-4o-mini)n/an/an/a14.01.0
Method rows are ours, on the pilot sets (83 LongMemEval-S questions, 100 per MAB bench, equal k). “not run” marks a cell we did not run. Published rows are copied from the papers’ tables, named in each label, and are not ours: MAB cells are SubEM accuracy in percent at the 262K context; LongMemEval cells are on the paper’s 0 to 1 scale. n/a means the source reports no such cell.

How to read thisHorizon’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.

7. What would change our mind

  • If the trained embedding scores below Dense and the tanh version scores below that, the geometry did not help anywhere, and the archive’s own April note was right.
  • If norm does not track generality, the picture in the middle of this page is not what the embedding learned, whatever it scores.
  • If the trained embedding matches Dense and the norm tracks generality, the geometry is real and not needed for these benchmarks. That is a finding, not a failure.

Sources

  1. Archived manifold tests · cosine against a tanh-projected Poincaré ball on synthetic clusters; the project’s own conclusion
    measured · internal · engram/archive/baleen_old/benchmarks/results/NOTES.md and manifold-compare.json · as of 2026-04-09
  2. Archived theory note · memory as a dynamical field; reads and writes both perturb it
    context · internal · engram/archive/baleen_old/docs/architecture/theory.md · as of 2026-04
  3. Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
    context · internal · engram/archive/baleen_old/benchmarks/SCORECARD.md · as of 2026-04
  4. context · public · as of 2017-05-22
  5. The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
    derived · internal · engram/bench/README.md, engram/bench/data/SOURCES.md · as of 2026-10-08
  6. Published rows · every number copied from the two papers’ tables, with the table named
    measured · internal · engram/bench/data/published.json · as of 2026-10-08
  7. context · public · as of 2026-06-28
  8. context · public · as of 2025-03-04
  9. Rebaseline plan · the ten experiments and their improved designs
    context · internal · engram/RESEARCH_PLAN.md · as of 2026-10-08

Experiment 6 of 10 in the Engram series. ← Bloodhound · Engram overview · Loom →

Horizon is one of ten experiments in Engram. Its rows come from one run on the pilot sets.