[home] [coding projects] [research projects] [research interests]
An active research problem about fixed-size persistent memory in sequence models.
Active research. I previously used this page as an implementation manual; current mechanism, source excerpts, equations, and benchmark details have been removed while the project is ongoing.
Attention and memory are adjacent but different operations. Attention decides what is useful for the current computation. Persistent memory decides what should remain available after the immediate context has moved on.
That distinction matters when a model has a short local window but must preserve a small amount of information over a much longer delay. Simply caching everything is one answer, but then state grows with context length.
What properties should a fixed-size persistent state have if it is expected to retain useful information over long delays while remaining strictly causal? The central tradeoffs are memory capacity, temporal resolution, stability, state size, and computational cost.
For this project the useful preliminaries are dynamical rather than software-specific. I want causality, fixed state, stability, and delayed retrieval defined before talking about any particular memory architecture.
These preliminaries are enough to state the public problem. The nonlinear state equations, gating rules, timescale construction, and current experiments remain private while the project is active.
I am comparing broad classes of persistent-state mechanisms under controlled memory budgets and asking which structural properties survive as retrieval delay increases. The exact state equations, update rules, ablations, and current results are intentionally withheld.
Last updated: September 14, 2026.