Glyph
A memory that carries its own manual, so a model that has never seen it can read it.
Glyph, replayed
MARCO? needed a page of instructions before four model families would agree on a word. Glyph tries to remove the page. The definition of the language is the first hundred or so tokens of the file, so whoever opens the file is taught by it.
1. The idea
Experiment 10 of 10 in the Engram series. ← Canopy · Engram overview
I want an agent’s memory to be something a different agent can pick up. Most memory is a private habit of the system that wrote it: vectors only one encoder understands, or summaries in one model’s manner. Change the model and the memory has to be rebuilt, or is quietly misread. That is the same problem as the one under MARCO?, an agent that cannot take its work to another platform, seen from the memory side.
Glyph’s bet is simple to say. Write each memory as one line of a compact language, and put the definition of the language at the top of the same file. The alphabet is the system prompt. A model that opens the file reads the definition first, then the memories, and has everything it needs.
Glyph · the header is the manual
A memory that carries its own manual is a memory that can change hands.
Decay is the other half of the original idea. Each line carries a half-life marker and the reader sees how old it is, so the model itself judges how much to trust it. No scheduler is involved. This run does not test that half: neither benchmark varies time in a way that isolates it, and one thing at a time is the point of the rebaseline.
2. Where it came from
The archived version was the namesake of the whole programme. It had a seed alphabet of nine frames (stable fact, completed event, open concern, plan, and so on), six half-life markers from an hour to forever, and a rule that a working set of seven lines is shown to the model at a time. The notation was written by Haiku and read by Sonnet. The paper behind it claimed about eight tokens per memory against about fifty for a sentence, and “cross-model fidelity: confirmed”. Later generations let the model rewrite its own language and runtime, and I have not carried that forward.
The audit’s verdict was that this was the cleanest new idea in the archive and the least tested. The token claim was never measured: the code counted whitespace-separated words. The cross-model result was a few hand-picked questions with no score. The attempt on MemoryAgentBench truncated each context to 1,500 characters, so most of the facts a question needed were never written; one run scored zero rows, which is a failed run and not a zero; and a later generation scored 4 of 100 under a home-made judge through an adapter that skipped the memory altogether. None of it measured whether one model can read another’s notation, which is the claim that matters.
3. Hypothesis
Stated before running, in three parts, because they can fail separately.
- Fidelity. Model B, reading a Glyph store written by model A with no primer, answers FactConsolidation questions as accurately as it does from the plain text of the same facts. The difference, paired, has a 95% interval that includes zero.
- The file does the teaching. B reading A’s store scores about as well as A reading its own, and B reading the store with its header removed scores clearly worse. If it does not, the header is not what makes the language readable.
- Tokens. Glyph stores each fact in fewer tokens than plain text, counted with a named tokenizer and with the header included.
My prior is that the third part is the weak one on this data. FactConsolidation facts are already short sentences, so there is not much to take out. The header costs about a hundred tokens once, which over the 455 facts of a 6K context is under a quarter of a token per fact, so even a small saving repays it. The question is whether there is a saving at all, and whether compression costs accuracy. I expect it to, a little.
4. Design
- The store
- The seed alphabet as a header, then one line per fact. Model A writes each line from the plain fact. For LongMemEval a session becomes a block of lines and retrieval returns session ids, as for every other method.
- The reader
- Model B, from a different family, receives the official FactConsolidation prompt with the retrieved lines in place of the facts. The prompt does not mention Glyph, its syntax, or the header. There is no primer.
- Retrieval
- The Dense scorer runs over the Glyph lines themselves, at the same k as every method. That is deliberate. If compression removes what a retriever needs, the score should say so.
- Controls
- A plain-text store of the same facts. The same store read by the model that wrote it. The same store with its header removed, so B sees lines and no definitions.
- Round trip
- B is also asked to restate each retrieved line in plain English. The restatement is scored on whether the entities and the relation survive.
- Models
- Writer and reader are named in the run record and held fixed across all items. One pair first. A second pair only after the first has been reported.
5. How it is tested
The measurement is the cross-model one. The answer score is the official SubEM on B’s reply, cut to its first ten tokens, as for every FactConsolidation run. Two measures do not fit the shared columns and are reported as text beside the table: round-trip fidelity, and tokens per fact for Glyph and for plain text with the header amortised across the facts. Each is given with its interval and the tokenizer is named.
MARCO? is the nearest thing to evidence we have. Four model families, handed the same 45 public statements and the same encoder, placed six situations in the same place 34 times in 36, against 4 in 36 for shuffled labels. That test shared a primer and a vocabulary. Glyph is a stricter one: the only thing the two models share is the file.
MARCO?’s rule, that a belief must never be written as if it were evidence, is a natural extension for a memory language: a line could say whether the fact was checked or only recalled. The first run leaves it out so that the test measures one thing.
The standing rules for every Engram experiment apply here unchanged.
- Fixed, seeded pilot sets: 100 LongMemEval_S questions (83 scored under the official skipping rules) and the 100 questions each of FactConsolidation single-hop and multi-hop at 6K. The seed is 20261008 and the sets are committed.
- Equal k. Every method, ours and the anchors alike, returns at most the same k ids, and the runner asserts it.
- No training or warm-up on scored data. Anything learned uses a seeded split disjoint from the pilot sets, and the code asserts the disjointness.
- Official metrics: recall_any@k and NDCG@k for retrieval, and the official SubEM with the official FactConsolidation prompt for answers.
- 95% confidence intervals on every number (Wilson for 0/1 metrics, bootstrap for NDCG), and paired tests against BM25 and Dense on the same questions (paired bootstrap and McNemar).
- Stochastic methods run on five seeds; the table reports the mean and the range.
6. Results
Glyph · results
| Method | LongMemEval-S recall@5 | LongMemEval-S NDCG@10 | MAB FC-MH 6K answer in top 10 | MAB FC-SH 262K SubEM | MAB FC-MH 262K SubEM |
|---|---|---|---|---|---|
| Glyph, read cold by a different model: 88.0 SubEM on FC-SH 6K, n=50 | n/a | n/a | n/a | n/a | n/a |
| Same facts in plain English (upper bound): 100.0 SubEM, n=50 | n/a | n/a | n/a | n/a | n/a |
| BM25 (ours) | 0.928 | 0.875 | 0.27 | 75.0 | 0.0 |
| Dense bge-small (ours) | 0.952 | 0.885 | 0.39 | 72.0 | 0.0 |
| BM25 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 48.0 | 3.0 |
| Text-Embed-3-Small (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| Mem0 (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 18.0 | 2.0 |
| MemGPT (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 28.0 | 3.0 |
| MIRIX (published, MAB Table 3, GPT-4o-mini) | n/a | n/a | n/a | 14.0 | 2.0 |
How to read thisGlyph’s rows are one run on the pilot sets with the new harness; no archived number is carried over. The LongMemEval columns are session retrieval (83 questions). The FC-MH 6K column is retrieval only: does any of the 10 returned facts hold the answer (100 questions; a 95% interval on a single cell is about ±0.1). The 262K SubEM columns are answer accuracy, which we ran for the BM25 and dense anchors only; the published rows sit at that context. “not run” means exactly that.
7. What would change our mind
- If B scores the same with the header removed, the header teaches nothing. Either the notation is guessable or the questions never needed it, and “self-describing” has not been shown.
- If A reading its own store scores far above B reading it, what looked like a language was one model’s habit. The memory did not change hands.
- If plain text matches or beats Glyph on accuracy and on tokens, there is no case for the notation on this data, and I will say so.
- If Glyph holds accuracy and saves tokens on FactConsolidation, the next step is to repeat it on LongMemEval and a second model pair before anyone should believe it.
Sources
- The archived notation · the seed alphabet (frames, half-life markers, the hot stack), written into the store itself
- The archived working paper · non-technical overview of the notation, with its claims and its examples
- Archived scorecard · the first round’s hand-written results table, with no raw logs behind it
- The Engram bench · one Memory interface, fixed pilot sets, official metrics and prompts
- Published rows · every number copied from the two papers’ tables, with the table named
- Rebaseline plan · the ten experiments and their improved designs
Experiment 10 of 10 in the Engram series. ← Canopy · Engram overview
Glyph is one of ten experiments in Engram. Its rows come from one run on the pilot sets.