Blog · Training Data · October 2, 2026

Synthetic Training Data Needs a Family Tree

A generated dataset's percentage tells us little about its history. To interpret a training result, we need to know what was retained, what was replaced, and whether the test shares ancestors with the training material.

The same label, different histories

Imagine two hypothetical teams reporting that their training data is mostly synthetic. One keeps its starting corpus and adds each generation of model outputs. The other repeatedly discards the previous training set and substitutes newly generated examples. Their summary labels might look alike. Their models have encountered different histories.

A reader asked to assess the results needs more than a synthetic percentage. Which examples survived? Which model generated the additions? Did training actually revisit the retained material? Was the reported improvement measured against a test that helped produce the new examples?

Our argument is that synthetic training data needs a family tree: a record of derivation and use, attached to a particular training run. This is a proposal for interpreting evidence, not a claim that documentation prevents model collapse. It narrows the broader concern explored in When the Training Set Starts Eating Itself to a practical question: what history must a reviewer reconstruct before believing the result?

Read the training protocol

Shumailov and colleagues' 2024 Nature paper demonstrates recursive degradation, including loss of low-probability parts of a distribution. Its language experiments repeatedly fine-tuned OPT-125m using generated text. A condition retaining some original data degraded less than one retaining none. The paper distinguishes finite experiments from theoretical behavior over indefinitely many generations. These are specific training settings, not observations of every deployed model.

Gerstgrasser and colleagues' accumulation study supplies important counterevidence. Retaining the initial corpus and successive generated datasets avoided progressively worsening error in their tested settings. Their linear-regression analysis bounds error independently of the number of iterations under its assumptions. The language experiments began with TinyStories, itself model-generated. Their baseline corpus therefore should not be casually described as human-authored ground truth. Accumulating data also meant more training steps per epoch as the corpus grew; the paper includes dataset-size ablations. Bounded error is not a promise of continual improvement.

Our reading is that the disagreement cannot be resolved by counting synthetic examples alone. A retained starting corpus, a fresh sample from an outside source, and another batch from the previous model are different interventions. So are training from a new initialization and continuing to update existing weights. A useful comparison identifies those choices before treating the findings as competing verdicts on synthetic data.

It also names what deteriorated. Average prediction error, coverage of unusual examples, and output diversity answer different questions. We would reserve a collapse claim for an identified pattern across training generations, with a stated measure and reference distribution. A disappointing score after one change is a reason to investigate; it does not identify recursion as the cause.

What belongs in the family tree

There is an established vocabulary for this task. The W3C PROV overview describes provenance through the entities, activities, and people involved in producing something, with support for representing derivation and exchanging records. It is a general provenance framework, not a synthetic-training quality certificate.

The following application is our proposal. Treat each dataset version and model checkpoint as an identifiable object. Connect them through activities: generation, rewriting, filtering, mixing, and training. A family tree is convenient language, but the actual structure is a graph. A batch may draw on several source collections, and the model producing it has a separate training ancestry.

Record both immediate parents and the boundary of knowledge. A generated answer might depend on a prompt, retrieved documents, and a teacher checkpoint. If the teacher's training data is undisclosed, mark that branch unknown. Calling the output synthetic does not fill the gap; naming the teacher does not establish everything it learned.

Keep origin separate from experimental role. An initial dataset can be the reference distribution for a study while itself containing generated text. Conversely, a newly collected document might describe an actual event while including AI-assisted editing. For a particular claim, specify what makes the reference suitable rather than relying on a universal division between real and fake.

The record also needs the operation performed at each generation. Did the team append examples, replace the previous set, retain only a selected subset, or bring in newly collected material? Record the generator version, generation settings, filter version, and accepted and rejected batch counts. Those details let a reviewer locate where two otherwise similar runs diverged.

Most importantly, distinguish storage from exposure. Suppose, hypothetically, that a team archives its starting corpus but draws every training batch from the newest generated material. The archive remains intact, yet the current optimization never encounters it. Our recommendation is to record sampling weights, tokens actually consumed from each source, and repeated passes. A retained file cannot explain a training result unless its role in training is visible.

This need not mean publishing every example. A restricted corpus can have a version identifier, a documented custodian, and a reproducible manifest available to an authorized reviewer. The public account should say which relationships were inspected and which remain assertions. Provenance is useful when it makes uncertainty legible as well as recording what is known.

The test has ancestors too

Yang and colleagues' study of rephrased benchmark samples found that changed wording could evade string-based overlap checks while distorting benchmark performance. They also identified contamination in the generated CodeAlpaca dataset. The experiments establish a weakness in those checks; they do not establish that every synthetic dataset contains benchmark material.

Our inference is that the evaluation requires its own ancestry record. A test item may have been excluded from the final training directory while a transformed relative entered through generation. Record whether evaluation material was available during prompt design, example creation, filtering, or model selection. An unknown teacher history should remain an uncertainty in this account.

Contamination and collapse must stay separate. Contamination concerns what a score can establish about performance on unseen material. Collapse concerns deterioration through a model-data feedback process. A rising benchmark score alone cannot establish that coverage survived recursion; falling coverage alone does not prove the benchmark was contaminated. The same run may require both investigations.

This boundary also separates the present proposal from the site's phantom-disclosure privacy audit. A lineage record does not measure private-information leakage. It can help locate inputs for an audit, but the performance claim and the privacy claim still need their own evidence.

A comparison someone can review

Return to the hypothetical teams. Before comparing scores, request their starting-corpus identifiers, generation history, actual sampling schedules, training budgets, and evaluation boundaries. Then ask them to state the conclusion at the same level of specificity: performance under this mixture and training procedure, measured against this reference.

Our proposed comparison would hold the evaluation fixed while separately examining retained history, new generated material, and newly collected observations. Where compute differs, report that difference alongside results. Include measures chosen to expose the losses that matter for the intended use, such as performance on a predeclared set of unusual but valid cases. These are design recommendations, not experiments conducted for this article.

A family tree will not decide whether a synthetic batch is good. It will let someone ask a sharper question about the result: what changed, what evidence remained available, and what alternative explanation is still plausible? That is enough to improve the conversation. Useful synthetic training should be judged by a reconstructable procedure and a credible evaluation, with the unknown branches left visible.

Sources

Sources consulted October 2, 2026. These studies establish results under their stated protocols; this article does not measure the condition of current production models.

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog