Blog · Science and Evidence · October 2, 2026

A Null Result Still Matters When AI Does the Searching

An AI-assisted search can produce a persuasive finding while leaving its unsuccessful tests out of view. Keeping those outcomes visible changes what readers can infer—and what the next experiment needs to establish.

The missing attempts

Imagine a research group asking an AI agent to improve a forecasting method. This is a hypothetical example. The agent tries different features, changes its evaluation window, abandons several approaches, and eventually finds a promising combination. The resulting paper contains working code, accurate numbers, and a convincing explanation. A reader can reproduce its chosen analysis. What the reader cannot see is how much searching preceded it.

That missing history matters even if nobody fabricated anything. The paper answers what happened along the selected route. Readers also need to know why that route became the one worth reporting. Our argument is that AI-assisted research should preserve the outcomes of substantive tests and make a clear transition from finding a promising question to testing it against fresh evidence.

The underlying problem predates research agents. In their 2014 study of publication bias, Annie Franco, Neil Malhotra, and Gabor Simonovits tracked 221 survey experiments from the TESS program. Strong results were 40 percentage points more likely to be published than null results and 60 points more likely to be written up. Much of the loss occurred before journal submission. This was a particular social-science cohort, not an estimate of bias in AI-assisted science.

Our inference is that automating manuscript production will not by itself solve selective visibility. A system can make writing cheap while still being directed to write only about its most attractive outcomes. The choice of what deserves a manuscript remains consequential even when producing that manuscript takes little effort.

What search selects

The AI Scientist-v2 report, version 1 from April 2025, describes agents that execute experiments, refine promising branches, and carry selected nodes into later stages. Its evaluation also involved researcher selection of initial ideas and submitted manuscripts. Yet its workshop-accepted manuscript reported negative findings, and was withdrawn afterward under a prior agreement. That outcome complicates any claim that automated science necessarily suppresses disappointing results. It was a limited workshop evaluation, not evidence of general scientific reliability.

Selection is necessary: no group can investigate every possible idea. The question is what readers are told about the selection that shaped a claim. Choosing a topic because it is interesting differs from choosing an analysis because its outcome looks favorable. Both choices may be reasonable during exploration; their evidential consequences differ.

This is where our focus departs from the site's essay on automated reanalysis. Recovering a published number checks an important property of the surviving result. It cannot, by itself, reveal which other tests never reached the paper. The research-idea funnel raises a neighboring question about which proposals get generated. Here the concern begins when those proposals meet data.

Record tested claims

We propose a compact record organized around questions that were actually tested. It should distinguish an idea rejected before execution, an experiment that failed technically, a completed analysis that remained inconclusive, and a completed analysis that challenged the prediction. A broken data loader is evidence about the workflow; it does not refute the scientific hypothesis.

For each substantive test, retain the question, data version, analysis or code, outcome, and reason for continuing or stopping. Connect revisions to the earlier attempt so that a changed outcome measure does not appear to have been the plan all along. Record which results were available when each decision was made. A short index linked to saved artifacts can make this usable without requiring readers to inspect every agent message.

The unit matters more than the volume. Counting every prompt as an experiment would inflate the apparent search. Counting only completed papers would hide it. Teams should explain their grouping rule and distinguish repeated runs of the same analysis from changes to the hypothesis. We recommend recording these decisions as work proceeds, because an attractive final narrative can make earlier uncertainty difficult to reconstruct.

This record would support interpretation, not automatically supply a statistical correction. Readers should be able to locate the evidence behind a selection decision without assuming that a count of attempts captures every dependency among them.

Protect fresh confirmation

Dwork and colleagues' 2015 research on adaptive data analysis explains why a testing dataset can itself become overfitted when feedback repeatedly guides subsequent choices. They develop controlled methods for holdout reuse under stated assumptions. The relevant lesson is that calling data a “holdout” does not preserve its independence when its results steer the search.

Our proposed workflow gives exploration room to operate, then marks a deliberate commitment. Before accessing confirmation outcomes, fix the selected claim, primary measure, analysis, exclusions, stopping rule, and criteria for interpreting the result. Timestamp the plan and disclose what the exploratory phase already revealed. Where several claims advance together, specify how their joint testing will be handled.

In the hypothetical forecasting project, the team could select a method using development data, freeze the comparison, and evaluate it on an untouched period appropriate to the intended claim. The test would still need to address how well that period represents future use. Holding data back does not answer every question about applicability.

If the team examines the new result, changes the method, and tries again, our recommendation is to document that return to exploration. A newly opened agent conversation is not a scientific reset if the new instructions encode what the team learned. Preregistration should clarify this boundary; it cannot erase earlier access to outcomes.

Commit before the outcome

The Center for Open Science's Registered Reports framework adds a publication commitment to advance planning. Reviewers assess the proposed question and methods before outcomes are known. In-principle acceptance depends on following the protocol and meeting quality requirements, rather than obtaining an attractive result. Additional exploratory analyses can be reported separately. This differs from simply posting a preregistration, which does not itself secure a journal's commitment.

There is evidence that the resulting literature looks different. Scheel, Schijen, and Lakens's 2021 comparison found support for the first hypothesis in 44% of 71 psychology Registered Reports, versus 96% of 152 standard reports. The Registered Reports sample reflected records available in November 2018. This was observational: hypotheses, researchers, and journals were not randomly assigned, and sampling language could affect the comparison. The gap does not isolate a causal effect of publication format.

For an AI-assisted project, our interpretation is practical: decide which selected questions merit a careful answer before knowing whether that answer will be exciting. A lab can reserve time and resources to finish and report such tests even without a participating journal. That local commitment is weaker than editorial acceptance, but it makes abandonment a decision to explain.

Make the null informative

A nonsignificant result does not establish that an effect is absent. Daniël Lakens's primer on equivalence testing explains how prespecified bounds can instead support a narrower claim: effects large enough to matter can be ruled out by an appropriate test. Insufficiently informative data may establish neither a meaningful difference nor equivalence.

Our recommendation is therefore to report what the failed prediction actually teaches. State the setting, estimate, uncertainty, and size of effect the design could meaningfully assess. Explain whether the study constrained the proposed explanation or simply left the question open. These distinctions make an unsuccessful test useful to someone deciding what to try next.

Return to the forecasting group. Its most valuable contribution might be a promising method, a well-supported limit on that method, or an honest account of unresolved uncertainty. Each can guide later work if the tested question and outcome remain accessible. A research system earns trust partly through the evidence it preserves when the preferred answer does not arrive.

Sources

Sources consulted October 2, 2026. Historical studies describe their own samples and systems, rather than current field-wide rates.

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog