What Would Count as Evidence That a Model Forgot?
A missing answer is an observation. Removing training influence is a claim about how a model was made. Understanding the comparison between them makes unlearning evidence more useful—and its limits easier to see.
The missing answer
A hypothetical publisher removes a withdrawn author's biography from a training collection. Before the change, its assistant supplied the author's birthplace. After an update, the assistant declines to answer. The publisher now has evidence of changed behavior. What else has it established?
That depends on the intervention. Removing the biography from a collection changes which records are present. Blocking an answer changes what reaches a reader. Changing a model's parameters changes the model itself, but the existence of a change does not establish that the biography's influence disappeared. These are different claims even when the user sees the same silence.
The distinction matters because a successful intervention can be oversold by its description. A refusal may satisfy a product requirement to stop displaying a particular answer. Calling that result complete forgetting adds a proposition the observation alone cannot support. Conversely, useful evidence of removal need not take the form of a model becoming unable to discuss an entire subject.
This essay's interpretation is that the decisive question comes before the score: what would the appropriate comparison model be? Without that comparison, an evaluator can reward silence, punish legitimate knowledge, or mistake general deterioration for selective removal.
Which model should it resemble?
Bourtoule and colleagues' Machine Unlearning formalizes an exact target: the distribution of models after removal should equal the distribution produced by training without the removed data. Training randomness matters; the requirement is not identical weights between two individual runs. The paper appeared at IEEE Security and Privacy in 2021.
Applied to the hypothetical biography, this comparison asks what the specified training process would produce without that record. It does not ask for a system incapable of stating the author's birthplace under every circumstance. Suppose another retained publication independently supplies the same fact. A model trained without the withdrawn biography could still learn it. That hypothetical case separates removing a record's contribution from prohibiting knowledge of a fact.
This also exposes a choice hidden inside the word “target.” Is the request about one document, every copy of its text, or everything concerning an author? Each defines a different collection to exclude. An evaluator cannot resolve that choice merely by selecting a more sophisticated score. The intended target must be explicit before the comparison has meaning.
Our practical inference is that a retraining reference needs a stated starting point. Repeating a later training stage without a document answers a narrower question if an earlier checkpoint already encountered it. A carefully constructed comparison can still be valuable; its description should identify which exposure it excludes. Otherwise, the reference quietly assumes the very absence it is meant to establish.
Evidence built into training
The same paper introduces SISA: separate models learn from isolated portions of the data, with saved intermediate states. Removal can then retrain the affected component from a state preceding the relevant data, excluding those data, before combining component predictions. This is a training design, rather than a retrofit that can simply be applied to any finished model. Its isolation requirement is central. The SISA construction supplies the argument for removal.
The lesson we draw is constructive. Evidence need not consist solely of unsuccessful attempts to elicit an answer. A system can be designed so that its training history supports a specific removal argument. That shifts some scrutiny toward whether the actual implementation followed the construction. A sound method on paper and a faithful run of that method are separate things to establish.
Guo and colleagues' Certified Data Removal from Machine Learning Models, published at ICML in 2020, offers a different guarantee: bounded statistical distinguishability from training without the removed examples. Their mechanism uses regularized linear models under specified mathematical assumptions, combining an approximate update with randomness that masks residual effects. Here, “certified” refers to a formal bound. It does not mean an unrestricted guarantee for an arbitrary language model.
These constructions are evidence against blanket pessimism. They also show why the adjective “approximate” is insufficient on its own. A claim might mean that a procedure satisfies a defined bound, or simply that measured behavior looks close to a reference. Readers need to know which meaning applies. Numerical closeness on a selected score does not automatically inherit a mathematical guarantee.
What a benchmark can establish
The 2024 TOFU paper makes the comparison tractable using fictitious author information introduced during fine-tuning. It compares unlearned models with a reference fine-tuned only on retained data and evaluates both forgetting and usefulness. The authors explicitly identify a limitation: a model producing nonsense can obtain a favorable forgetting score. Their separate utility measurements help expose that failure. They also limit the task to fine-tuning exposure, rather than removal from pretraining.
Our reading is that two judgments must remain distinct: whether the measured difference was small, and whether the measurement could reveal the differences that matter. A test that fails to distinguish models supplies evidence within its scope. It does not establish equality across all possible behavior. A handful of unanswered prompts gives an especially narrow view.
Usefulness belongs beside the forgetting result because selective change is the ambition. If the hypothetical publisher's assistant stops answering every literary question, the withdrawn biography will also stop appearing. That outcome could satisfy a crude absence measure while defeating the reason for maintaining an assistant. Evaluating retained work makes the cost visible instead of allowing collapse to masquerade as precision.
Grimes and colleagues' 2024 Gone but Not Forgotten examines evaluation weaknesses in computer vision. It argues for attention to worst-case outcomes, information revealed across model updates, and performance through repeated removal requests. The study offers preliminary evidence in its experimental setting; it is not a measurement of deployed language models in October 2026.
The broader implication is a limit on transferring conclusions. An average score does not describe every removed record, and a result after one update does not describe a long sequence. Neither observation makes the original experiment worthless. It identifies the additional claim that would need evidence before a provider extends its promise.
A claim with usable boundaries
We propose reading an unlearning report as an argument with four connected parts: its target, its comparison, its evidence, and the useful behavior it preserves. This is an editorial framework for interpreting claims, not a new certification standard. The parts should agree. A document-level target cannot silently become a claim about erasing a person's entire presence.
For example, a hypothetical report might identify a removed collection and a particular checkpoint, describe a retained-data reference, and report behavioral comparisons alongside the assistant's continued performance on unrelated editorial tasks. If no formal bound applies, the report should describe its result as empirical. If a bound does apply, the assumptions and covered training procedure should accompany it.
That description leaves room for meaningful success. A provider may establish that a specified procedure excludes particular training examples, or that a revised model closely matches a defined reference on measured behavior. Either is more informative than an unqualified promise of forgetting. The reader can see what has been established and decide whether it answers the actual request.
For an institution concerned with public memory, this precision has a further value. Withdrawal should be describable without pretending that knowledge is a single object with a delete button. A responsible account names the contribution removed, the evidence supporting removal, and the uncertainty that remains. That gives both remembering and forgetting an accountable history.
Sources
- Bourtoule et al., Machine Unlearning. IEEE Security and Privacy, 2021; arXiv version 3, December 15, 2020.
- Guo et al., Certified Data Removal from Machine Learning Models. ICML, 2020.
- Maini et al., TOFU: A Task of Fictitious Unlearning for LLMs. January 11, 2024.
- Grimes et al., Gone but Not Forgotten: Improved Benchmarks for Machine Unlearning. May 29, 2024.
Primary sources and their current records checked October 2, 2026. Research dates above identify the studies; they are not claims about current product performance.
Related reading
- The Unlearning Claim Becomes the Localization Test
- The Training Opt-Out Becomes the Consent Interface
- Machine Unlearning
Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.