Blog · Learning and Assessment · October 2, 2026

What Students Can Do After the AI Help Is Removed

A completed assignment shows what a student can produce with the available help. To judge what an AI tutor has taught, separate immediate independence, delayed retention, and the ability to use knowledge on a changed problem.

The finished worksheet

Imagine a student completing an algebra worksheet with an AI tutor. Every answer is correct. The following lesson, with the chat closed, the student cannot decide how to begin a similar problem. This is a hypothetical classroom scene. It exposes an ambiguity in the worksheet: the finished page records a successful collaboration, but leaves the student's independent capability unresolved.

Closing the chat and testing immediately would clarify part of that ambiguity. Waiting before testing would clarify another part. Changing the problem would ask a further question. Our argument is that these assessment choices should remain separate whenever a school describes an AI tool as improving learning. An independent answer today, a remembered method later, and a useful response in a new situation are different achievements.

What removing help establishes

In Bastani and colleagues' high-school mathematics trial, both AI conditions improved assisted practice performance. On the subsequent unassisted exam, the unrestricted GPT Base condition scored 17% below the control group, a relative reduction. The guarded GPT Tutor condition had no statistically significant difference from control. Crucially, instruction, practice, and examination were contiguous parts of each session; exam problems closely resembled practice problems. Published in 2025, the study supplies evidence about learning after assistance is withdrawn, rather than a measurement of retention weeks later.

Positive evidence also deserves its actual scope. Kestin and colleagues' 2025 physics study found higher post-lesson scores with a carefully structured AI tutor than with classroom active learning. Its crossover design covered two lessons, and the tests were written by a team member separate from those designing the tutor or teaching the lessons. This is useful evidence for a particular instructional approach. The reported post-lesson comparison does not establish delayed retention or broad transfer beyond the assessed material.

Our reading is that these findings make a universal verdict on AI tutoring less useful than a precise description of the intervention and outcome. They differ in subjects, students, instructional arrangements, and comparisons. They are not a contest in which one result cancels the other. A successful tutor design can merit further use while still needing a later test of what persists.

The distinction also protects useful assistance from an unfair standard. If a student's objective is to finish an accessible practice activity with support, assisted success matters. If the objective is to solve that class of problem independently, the assessment must give the student that responsibility. The educational promise determines which outcome should carry the headline.

Time and novelty are different tests

The importance of delay predates chatbots. In Roediger and Karpicke's 2006 experiments, students learned prose passages through restudy or retrieval tests. Restudy performed better on a test after five minutes; retrieval practice performed better on delayed tests, including one week later. These were memory experiments without AI. They demonstrate that the timing of measurement can change which learning condition looks stronger, not that every AI intervention will show the same reversal.

Andrew Butler's 2010 experiments separately examined transfer. After studying passages, participants practiced through repeated tests or restudy. Final assessments one week later included new inferential questions, within or across knowledge domains. Testing produced better transfer than restudy in those experiments. Feedback accompanied the practice tests, so the results do not isolate retrieval from feedback, and they do not establish an AI tutoring effect.

For our proposed assessment approach, delay and novelty are independent choices. A delayed repetition asks whether something remains available. An unfamiliar problem asks whether the learner can recognize where knowledge applies. Either can be made harder without improving the other. A month-old question copied from practice may be a retention test; a new question asked immediately may test transfer without demonstrating durability.

Consider a hypothetical lesson on proportional reasoning. Substituting new numbers into the same layout checks a fairly close application. Presenting a graph and asking whether the relationship is proportional changes the representation. Mixing proportional and nonproportional cases asks the learner to choose the method. Those changes should be named, because a claim of “transfer” becomes more informative when readers know what actually changed.

An assessment sequence worth trying

The following is our design proposal, not a protocol validated by the cited studies. Start by naming a narrow capability: for example, choosing and explaining an appropriate method for a family of algebra problems. Write the assessment tasks before adapting the tutoring lesson. Decide what counts as success, including how partial reasoning will be credited, so an attractive result cannot quietly redefine the learning goal.

During practice, record assisted performance separately. Immediately afterward, give a short independent assessment using comparable problems. Later, administer a second independent assessment with both comparable and changed problems. A week is a practical hypothetical interval for a classroom pilot, not a scientifically established optimum. Choose a delay that matches when the course expects students to use the knowledge again.

Describe independence precisely. Remove the tutor and its conversation history, while preserving ordinary accommodations and whatever reference materials the target skill legitimately permits. If using a formula sheet is part of the intended competence, banning it would measure an additional memory demand. If choosing a formula is the intended competence, supplying the appropriate formula beside every question would remove the decision being assessed.

Keep comparable and changed items distinguishable in the results. Score the choice of method and explanation as well as the final answer where those are learning goals. Have assessors work without knowing which instructional condition produced the response when feasible. Review the changed tasks for accidental demands on vocabulary or background knowledge that the lesson never intended to teach.

There is a complication: testing can itself improve subsequent retention, as Roediger and Karpicke demonstrated. An immediate assessment is therefore also part of the learning experience. Our recommendation is to give comparison groups the same assessment schedule and feedback opportunities. That estimates the effect of the instructional package under that schedule; isolating the contribution of the immediate test would require an additional comparison.

A fair local trial should also record intervening instruction, extra practice, and missing follow-up results. Random assignment, where feasible, strengthens a comparison with the existing teaching approach. Merely comparing volunteers who enjoy the tutor with students who avoid it leaves other explanations open. Small pilots can expose practical problems, but uncertain estimates should remain uncertain rather than becoming a confident promise.

Read the pattern before choosing the headline

Our interpretation would depend on the pattern. Better assisted work with unchanged independent results supports a claim about supported performance. Better immediate independent results with no delayed advantage supports a narrower learning claim than durable improvement. Better delayed results on familiar tasks with weak performance on changed tasks points toward a capability that remains tied to the practiced format.

These patterns suggest different next steps: revise how practice distributes responsibility, add opportunities to revisit material, or teach students to recognize when a method applies. They do not by themselves diagnose the cause. A difficult transfer item might expose a teaching gap, an unfair question, or insufficient prior knowledge. Examine student reasoning before assigning the failure to the tool.

For a school choosing an AI tutor, the useful promise is specific: students taught under these conditions could later perform these tasks with this level of permitted help. That statement leaves room for valuable tutoring and for further improvement. It also puts the learner's subsequent capability at the center of the decision, where the completed worksheet alone could never put it.

Sources

Primary texts checked October 2, 2026. The older experiments inform assessment design; they are not evaluations of current AI products.

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog