An AI Memory Needs a Way to Become Out of Date
A preference, a provisional plan and a dated observation should not remain active for the same reasons. Useful assistant memory needs rules for when a remembered statement still applies, including when the honest answer is that its present status is unknown.
The plan that keeps returning
Consider a hypothetical assistant that remembers its user was thinking about joining a pottery class. Weeks later, every suggested weekend includes pottery. The assistant retrieved the right conversation. Its mistake was allowing a tentative intention to become a standing description of the person.
The user might have abandoned the idea, attended once, or simply postponed deciding. Silence establishes none of these possibilities. Yet treating the plan as indefinitely current also makes a claim without evidence. A useful memory must be able to preserve what someone said while reducing its influence on what the assistant proposes next.
LongMemEval makes relevant distinctions in its evaluation: information extraction, reasoning across sessions, temporal reasoning, knowledge updates and abstention are separate abilities. Its histories include model-generated evidence conversations screened and edited by people. This is a constructed evaluation, not a longitudinal observation of ordinary users. The paper also reports that good retrieval alone does not guarantee correct use of the retrieved material.
The design proposal developed here concerns that gap between finding a statement and deciding its present role. Earlier essays on stale-fact ledgers and timestamps lost in summaries address storage and preservation. The additional question is what kind of continuation the statement deserves.
Different kinds of continuation
Our proposed starting point separates preferences, plans and dated observations. These categories are imperfect, but they invite different questions before an assistant uses a memory.
A preference can provide a continuing default. In a hypothetical case, someone asks for concise answers. Requiring them to renew that preference every month defeats the purpose of remembering it. But the default needs a scope: asking for a detailed explanation on one occasion should not silently rewrite their general preference. Explicit changes should override defaults; local exceptions should remain local.
A plan has a different shape. Someone considering a pottery class has expressed an intention with uncertain commitment. Registering for the class would be stronger evidence, but even registration would not establish lifelong enthusiasm. A plan should carry whatever milestone or end condition the conversation actually supplies. Once its window passes, the assistant should stop treating it as an active prompt for action unless renewed. Its eventual outcome can remain unknown.
A dated observation is different again. Suppose, hypothetically, a user says the neighborhood studio has evening sessions this season. That observation can remain useful historical information while becoming insufficient evidence for next season's schedule. The appropriate response may be to check the studio's current information when a schedule is needed. No change in the user's preferences is implied.
Expiry therefore needs a precise object. The proposal is to expire a statement's eligibility to guide a particular present decision. That need not declare the statement false, erase its historical meaning, or invent a replacement. Whether to retain the underlying record is a separate choice, including the user's ability to remove it.
When learned, when applicable
Established database designs already distinguish time dimensions. MariaDB's bitemporal tables combine system versioning with application-time periods. Its application-time documentation describes intervals bounded by temporal columns. Those facilities supply representation mechanisms; they do not decide what a person's changing plans mean.
For assistant memory, the distinction matters whenever news arrives before or after a change. In a hypothetical conversation, a user says on Friday that their new working hours start Monday. The statement is newly learned on Friday, but the old schedule still applies over the weekend. Another user may report a change several days after it happened. Sorting memories by message date alone cannot handle both cases.
The memory should preserve the stated effective period and avoid supplying a precise boundary when none was given. “For a while” can remain an uncertain duration. Turning it into an invented expiry date would make the database tidy at the expense of the user's meaning.
TimeBench, published at ACL in 2024, evaluates temporal reasoning across symbolic, commonsense and event-based tasks. Its authors identify limitations including prompt-based evaluation and possible training-data contamination. It helps explain why interpreting temporal expressions deserves separate testing; it does not validate any particular expiry policy or establish how today's assistants perform.
A correction has a scope
A newer message need not replace an older preference. Consider another hypothetical user who normally wants morning meetings but requests an afternoon meeting while traveling. The relevant distinction is the trip's boundary. If the assistant promotes the exception to a permanent preference, it will seem forgetful immediately after making an update.
Correction should therefore identify what changed: the general default, one event, a temporary interval, or an earlier misunderstanding. A user saying they never wanted pottery lessons is correcting the interpretation. A user saying they no longer want them is reporting a change. Both should stop suggestions, but they describe different histories.
The W3C's PROV Data Model provides concepts for attribution, derivation and revision. These can describe where a stored claim came from and how a later record relates to it. Provenance is useful evidence about a record's origin; it does not by itself establish that the record still applies.
The practical interface should make the consequence of correction legible. A hypothetical acknowledgement could say that afternoon meetings will be preferred during the trip, with mornings remaining the usual default. That gives the user something specific to correct. A generic promise to remember leaves the scope hidden.
The cost of asking again
There is a strong objection to aggressive expiry: an assistant that constantly asks whether old preferences remain true transfers its maintenance work to the user. It can also weaken useful continuity. Persistent memory is valuable partly because settled choices need not be repeated.
Our proposed response is selective renewal at the point of use. A mild preference can guide a reversible suggestion without interruption. An unresolved plan should not repeatedly generate unsolicited nudges. A dated schedule should be checked when someone actually needs to rely on it. The question is whether renewed information would materially change the response.
Uncertainty should also change the wording. The assistant can describe a previously mentioned interest as previously mentioned, without presenting it as a current commitment. That small distinction allows useful continuity without requiring either endless confirmation or unwarranted certainty.
Test continuity as well as change
A proposed evaluation should include paired hypothetical histories. In one, a preference changes; in another, it stays stable while time passes. Add a temporary exception, an abandoned plan with no known outcome, and a future change announced early. Ask both present-tense and historical questions.
Score unnecessary expiry alongside stale persistence. Otherwise a system could appear successful by discarding every old statement or asking for confirmation every time. Also inspect whether it invents completion when a deadline passes. These are suggested tests, not experiments performed for this essay.
The cited research does not establish a universal duration after which personal information becomes unreliable. A useful assistant would instead preserve the distinction between remembering a statement, accepting it as currently applicable, and needing more evidence. People should be able to change their plans without having every passing intention follow them into future conversations.
Sources
- Di Wu and colleagues, LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory, version 2, March 4, 2025.
- Zheng Chu and colleagues, TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models, ACL 2024.
- MariaDB, Bitemporal Tables and Application-Time Periods, undated documentation.
- W3C, PROV-DM: The PROV Data Model, Recommendation, April 30, 2013.
Sources consulted October 2, 2026.
Related reading
- The Stale Fact Becomes the Memory Ledger
- The Memory Summary Loses the Clock
- The Memory Conflict Becomes the Write Transaction
Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.