01Abstract
Long-running agent memory systems commonly treat “importance” as a single quantity. We argue it conflates three distinct signals.
A scoring formula written in 2023 for a simulation demo—relevance plus recency plus importance—moved into production systems unexamined, letting the three jobs of “importance” (sedimentation, ranking, retirement—later separated as salience, ranking, and metabolism) strangle one another inside a single signal.
Our position rests on a service chain: memory serves the LLM; the LLM serves the human—and memory must be memory: self-describing and self-metabolizing, its salience and retirement decided by data and lifecycle rather than one-shot human-delegated verdicts. Skip the middle link and even the most elaborate memory system drifts toward RAG.
Since March 2026, over about seven months of production telemetry on a single conversational system, we argue that “importance” is at least three orthogonal signals that are best kept physically separated. We document three production failures of conflation, contribute a signal taxonomy with four invariants (I0–I3), seven design laws, and two same-day audit tools.
No benchmark supremacy is claimed; all evidence comes from our own failures, of which we have plenty.
02Three production failures
03Figures
04Two audits you can run today
The Silence Test. Suspend all retrieval for T days (or replay history offline while intercepting writes), then recompute your “importance / quality” metrics. If the numbers do not move, you measured data; if they drift, you measured behavior.
The RAG Test. Fix, in advance, a workload and an outcome criterion; replace your “memory layer” with a good-enough search engine. If behavior shows no material difference on that criterion, the layer was retrieval only—not memory.
05One honest boundary (and a correction)
Silencing retrieval shows zero drift because salience is computed from structural edges only; retrieval counters never enter the formula. Salience and access do correlate weakly (Pearson 0.24) — but that is a shared-structural-cause artifact, not a write path, and invariant I1 constrains the write path, not correlation.
We also corrected an earlier claim: the 0.999 quality metric was structure-blind (a ratio over link types, never counting real structure), not “behavior-contaminated”. The corrected scope-aware metric reads 0.101 on the same store.
06Contents
- Introduction: for whom does memory serve — the legacy of RAG
- Related work: three heirs and a critical wave
- Method: signal definitions, invariants, and a feedback analysis
- Case study: production failure evidence (single system, longitudinal)
- Design laws
- Tools: two audits you can run today
- Discussion: limitations, validity threats, and why no benchmark
- Conclusion: not sublimation, but regression