Back

TinyMem: learning fixed-size memory

I started TinyMem with a question: could a model learn to keep the useful parts of a history in a few hundred bytes? I built a writer that updates a small state as records arrive, then a reader that answers questions from that state. The writer never gets to see the question in advance.

I studied two related problems: whether repeated updates preserve earlier facts, and whether learned memory uses its storage budget more effectively than compact text. Controlled experiments showed that explicit state supervision could produce strong recall, but answer-only training remained unreliable.

In the final comparison on bAbI tasks 1–5, learned memory reached roughly 21% accuracy at 64, 256, and 1,024 bytes. Dictionary-coded text reached 99.7% at 64 bytes on clean inputs. All methods performed poorly under heavy distractors. That was not the result I was aiming for. The useful part was working out which failures came from the writer, which came from the reader, and what I could actually conclude from either.

I drew on recurrent memory models such as RMT[1] and ARMT.[2] ICAE[3] and Compressed Context Memory[4] informed the encoder–reader design review. These systems use different representations and storage accounting; TinyMem is not a replication of their architectures.

Writing and reading the state

The writer combines features from the current record with the previous memory state. After each update, the source text is discarded. At read time, a learned projection converts the saved state into prefix vectors for the language model; the question is introduced only at this stage.

Records are shown for context; only the state persists. The state pattern and answer illustrate the intended behavior, not a model prediction.

The later studies used Qwen3-1.7B[5] with frozen base weights. Some readers had trained rank-eight LoRA adapters. Those shared weights were separate from the per-history storage budget: the bytes that survived from one statement to the next.

From gated slots to selective updates

I first used gated slots, where a learned gate blends a candidate value into the old memory. The next design was informed by work on gated delta networks:[6] choose an address, read what is there, and write a correction. The final comparison used attention-based slots, quantized to int8 after every write.

Mt=Mt−1+kt⊗δtδt=vt−v^t\begin{aligned} M_t &= M_{t-1} + k_t \otimes \delta_t \\ \delta_t &= v_t - \hat{v}_t \end{aligned}

This is the basic update, before gating: Mt−1M_{t-1} is the old memory, ktk_t is the key, and vtv_t is the proposed value. The correction δt\delta_t subtracts the value already stored at that key, v^t\hat{v}_t. The outer product writes that correction into memory.

Overlapping keys were one possible source of write interference. The controls below test key functions, without isolating overlap from their other properties.

Reader competence

My first byte-level model, trained from scratch, scored 0 out of 500 on LongMemEval. It also failed an answer-copy control with the answer provided in its input. I could not blame the memory when the decoder could not handle that simpler task.

I therefore moved to simpler question-answering tasks and controlled readouts. The early byte-level model and the later Qwen-based system are separate experimental stages; their scores are not directly comparable.

Interference from repeated writes

A controlled test used four independent binary facts. After the writer encoded them, it received eight truthful repetitions of one fact. I measured recall of the three unmentioned facts before and after those writes.

This was the failure I found most useful to study: nothing new needed to fit, and no correct answer had changed. Yet recall fell in all three confirmation seeds, under both uniform and correction-weighted training.

Recall after repeated writes
-20.4 pp

Mean change in recall of the three facts
that were not repeated.

BeforeAfter eight repeatsAccuracy (%)
050100BeforeAfter eight repeats
Seed 202792.6 → 60.7%Seed 202882.4 → 62.9%Seed 202966.8 → 57.0%
Exact values by seed
Accuracy (%)
SeedBeforeAfter eight repeats
202792.5860.68
202882.4262.89
202966.8057.03

Repeating one true fact reduced recall of the other three by 20.4 points.

Three writer seeds, four binary facts, 66 bytes. The ground truth stays fixed; only one fact is repeated. Arms are paired within each seed. Lower recall does not prove that all information was erased.

A fixed linear probe recovered some facts the language model missed. Its recall also declined after repetition. This shows reduced recoverability under both readouts, without establishing that the information was irreversibly erased.

Initial recall versus retention

A subsequent delta-rule comparison used 66- and 258-byte states. Several models remained near chance and showed little dependence on their memory. Small before-and-after changes in these runs did not indicate good retention; the initial recall was already poor.

I needed a way to tell these failures apart, so I tested the components separately: reading a correct state, learning to produce that state, and retaining it through unrelated writes.

Testing the reader in isolation

I replaced the learned writer with an exact parser. Each of three fresh bridge-and-reader pairs scored 14,336/14,336 recorded answers, or 43,008 across all three readers. These records include shared-prefix duplicates and paired wording variants, rather than independent histories. Swapping in another story’s memory changed the answers to match that memory. These controls established that the reader could use the supplied state.

Stabilizing the writer

Training the writer against exact target states exposed a failure in some runs: gates or hidden activations approached zero. Adjusting write strength alone did not resolve it. Affine-free LayerNorm[7] before the hidden activation improved recall across the paired experiment, without adding parameters or storage bytes.

Effect of hidden normalization
258 B · 12 writer seeds
69.5 → 92.6%

Recall after unrelated repeats.

ControlHidden LayerNormAccuracy (%)
05010041014101: control 80.99%, LayerNorm 94.47%41024102: control 90.00%, LayerNorm 96.70%41034103: control 51.17%, LayerNorm 92.04%41044104: control 86.22%, LayerNorm 88.95%41054105: control 51.19%, LayerNorm 94.05%41064106: control 51.39%, LayerNorm 96.05%41074107: control 86.91%, LayerNorm 91.30%41084108: control 82.03%, LayerNorm 90.47%41094109: control 64.06%, LayerNorm 96.40%41104110: control 51.17%, LayerNorm 94.62%41114111: control 51.50%, LayerNorm 91.67%41124112: control 86.91%, LayerNorm 84.46%
Exact values by seed
Accuracy (%)
SeedControlHidden LayerNorm
410180.9994.47
410290.0096.70
410351.1792.04
410486.2288.95
410551.1994.05
410651.3996.05
410786.9191.30
410882.0390.47
410964.0696.40
411051.1794.62
411151.5091.67
411286.9184.46

LayerNorm improves the initial state, while retention still needs to be measured separately.

State-supervised delta writer, 12 seeds. Orange bars show controls and blue bars add hidden LayerNorm. Each bar averages three readers and two wording groups. Most of the gain came from better initial writing. Seed 4112 gets worse and stays in the result.

The catch was that most of the improvement came from learning a better state initially. It did not mean I had solved forgetting. The result also depended on exact state targets, a stronger training signal than answers alone.

Answer-only training

On the official bAbI QA1 test, state-supervised writers reached 98.2% mean answer accuracy, close to 98.3% with oracle memory. With answer-only writer training, the same architecture performed poorly: development accuracy averaged 29.1%. Those are different evaluation splits, so they should not be read as a matched test-set comparison.

The diagnostics found useful factual gradients, but also much more interference between unrelated writes. I then held the key function fixed and trained a fresh value branch. Keys from a successful state-supervised writer raised mean development accuracy from 31.8% to 59.0%, compared with keys from a failed answer-only writer.

31.8 → 59.0%

Overall development accuracy

20.6 → 58.5%

Earlier facts after unrelated writes

Six of the twelve pairs using successful keys still scored below 30%. The intervention improved mean accuracy but did not make value learning reliable. Because the reader and these keys had prior state supervision, this experiment was not a fully answer-only system.

I still needed to answer the original question: was any of this better than keeping a short piece of text? For the final study, I gave learned memory and explicit storage the same persistent byte budgets.

Matching the storage budget

It tested 64, 256, and 1,024 bytes across bAbI tasks 1–5. The learned writer and read adapter trained jointly from answer tokens. Every write was quantized, packed, and later restored from real bytes. Fixed-level scalar quantization[8] informed the choice to avoid a learned codebook; the int8 implementation is specific to this study. The alternatives kept compressed recent text, compressed text selected for lexical diversity, or a dictionary-coded text log.

Accuracy by storage budget
0255075100Accuracy (%)64 B256 B1,024 BDictionary recent text: 99.75%, seed range 99.64–99.82%Dictionary recent text: 99.75%, seed range 99.74–99.76%Dictionary recent text: 99.75%, seed range 99.74–99.76%Compressed diverse text: 62.32%, seed range 62.20–62.52%Compressed diverse text: 99.32%, seed range 99.32–99.32%Compressed diverse text: 99.75%, seed range 99.74–99.76%Compressed recent text: 66.99%, seed range 66.78–67.28%Compressed recent text: 99.75%, seed range 99.64–99.82%Compressed recent text: 99.75%, seed range 99.74–99.76%Learned int8 slots: 20.97%, seed range 20.90–21.02%Learned int8 slots: 20.73%, seed range 20.62–20.94%Learned int8 slots: 21.75%, seed range 21.04–22.50%
Learned int8 slots21.0%Compressed recent text67.0%Compressed diverse text62.3%Dictionary recent text99.7%

At 64 bytes, explicit text remains the strongest baseline.

Values and seed ranges
64 bytes · clean · accuracy (%)
MethodMeanMin–max
Learned int8 slots20.9720.90–21.02
Compressed recent text66.9966.78–67.28
Compressed diverse text62.3262.20–62.52
Dictionary recent text99.7599.64–99.82
5,000 official bAbI questions, tasks 1–5. Means over three seeds. Vertical marks show seed ranges, not confidence intervals. The byte axis is logarithmic. Shared weights and dictionaries are outside the per-history budget.

What the memory retained

The learned system stayed around 21% at every budget. Zero-memory controls averaged 17.2%, and other-story controls 16.6%. The repeated task-level scores and these controls indicated that the reader had mostly learned answer frequencies rather than a useful dependence on the stored state.

On clean inputs, the dictionary store reached 99.7% at 64 bytes, and compressed recent text reached 99.7% at 256 bytes. Under heavy distractors, every method fell to roughly 21–27%. The advantage of explicit storage was therefore specific to the clean setting.

Transfer to longer histories

On the known BABILong 1k transfer set, learned memory stayed near 20%. Compressed recent text reached 34.5%, 50.1%, and 88.4% as the budget increased. This was transfer to a benchmark the project had already encountered, not an untouched discovery set.

Evaluation protocol

All 5,000 official test questions were scored from the fixed final checkpoints. Training used three epochs and 16,875 steps per fit. There were nine learned fits, covering three budgets and three seeds, and three text-reader fits. No seed was removed for a poor score.

Text readers shared adaptation across budgets and policies, while each learned budget had its own adapted reader. The dictionary was built from training records. Shared dictionary bytes and model parameters were reported separately from the persistent bytes for each history.

Heavy noise added four distractor records after each four source records and a 16-record delay before reading. The chart’s ranges describe the three observed seeds, not uncertainty over all possible training runs.

Two evaluation-only code amendments were recorded. One restored the frozen-reader state after a feature-encoding context; the other repaired numerical error bars and report provenance checks. They reused the fitted checkpoints and did not change training. Failed partial outputs were retained.

I started TinyMem hoping that a learned state could make better use of a small storage budget than a short text history. In the systems I tested, it did not. Compact text was much stronger on clean inputs, while heavy distractors exposed weaknesses in every method.

The experiments changed how I think about that result. Reading a correct state, learning to write one, and preserving it through later updates are separate problems. Explicit state supervision gave strong recall, but that success did not carry over to reliable learning from answers alone. It also did not establish a compression advantage: four binary facts fit in one byte, however well a larger learned state recalls them.

My main takeaway is that a convincing memory system needs more than a good final answer score. It needs controls that show what was written, what remains recoverable, and whether a simpler store would do better. These experiments concern particular architectures, tasks, and training recipes, not a general limit on learned memory. They gave me a clearer account of where this approach failed, and of what a stronger result would have to demonstrate.

References

  1. [1]
    Recurrent Memory TransformerBulatov, Kuratov, and Burtsev (2022).
  2. [2]
  3. [3]
  4. [4]
  5. [5]
    Qwen3-1.7B model cardQwen Team (2025).
  6. [6]
    Gated Delta Networks: Improving Mamba2 with Delta RuleYang, Kautz, and Hatamizadeh (2024).
  7. [7]
    Layer NormalizationBa, Kiros, and Hinton (2016).
  8. [8]