MMLA: Bounded Editable Memory That Actually Decides What Deserves Authority
Long context lets models replay history. It does not decide which completed observations deserve to become authoritative state.
An agent running for months will face obsolete values, version conflicts, protected facts, and weak evidence. Simply stuffing more tokens into the window increases cost without solving the real problem: which past observations should update persistent memory, what they should overwrite, and when the system should refuse to write at all.
MMLA (Memory Management and Learning Architecture) formalizes a different answer.
We place a bounded resident memory between the transient context window and slow weight updates. Completed local segments are eventized. For each event a target-conditioned constructor proposes semantic content; a trusted assembler produces a complete versioned row (with tenant, ACL, version, protection, provenance). Deployment then makes a hard decision: atomic overwrite of an entire row, or NULL. Training may use realized futures to price actions; deployment remains strictly causal and future-blind.
Results that matter to researchers
Lifecycle control is exact. On a controlled versioned-variable suite deliberately designed so that last-mention retrieval is wrong by construction, the full lifecycle path scores 1.000 exact match on 300/300 held-out records across three seeds. Lexical baselines sit at 0.333. Generation with a frozen Llama-3.1-8B reaches 300/300 correct answers using only 146 prompt tokens versus 172/300 for full-context reading at 729 tokens.
Selection + calibrated sparse fallback beats dense retrieval while using far less evidence. On held-out HotpotQA (natural length and 8.2k-word conditions with answer-excluded distractors), a learned two-hop selector with a bounded 32-passage cache + archive fallback improves over a budget-matched dense baseline by 5.5–16.6 F1 and over BM25 by 4.0–6.2 F1. It reaches 102–116% of full-context F1 while consuming roughly 10% of the evidence words. The cache itself costs only 0–2.9 F1.
Typed relational transport is perfect under stress. Anchor–filler relational binding yields 240/240 whole-record exact matches per seed. Three same-checkpoint controls (candidate-content, unsigned-distance, type-erased) score 0/240.
All backbone LLMs remain frozen. We report the original Llama budget-gate failure and the subsequent Qwen adjudication explicitly. No leaderboard claims. No systems-superiority claims. Joint training, faithful architecture baselines, and full quality-cost measurements remain future work.
Why this is different
Most long-context and memory-augmented systems optimize for more recall or cheaper storage. MMLA optimizes for authority management: incorporation, overwrite, preservation, deletion/tombstoning, and abstention under hard capacity and security constraints. The resident state is bounded, explicit, versioned, and editable by construction.
The integration loop is not yet complete—semantic-interface readers on our frozen V28 checkpoint currently fail qualification—so predictive overwrite remains gated. We publish the negative results alongside the positive ones.
Paper: https://t.co/n029mTbbAR Code + full technical report: https://t.co/Il1O7Nn3e0
If you care about agents that must remain coherent for months rather than minutes, about memory that can be audited and budgeted rather than merely cached, or about architectures that separate training-time future supervision from deployment-time causality, take a look.
The past should shape the future. It should not simply be replayed.