An adult learner is refreshing algebra before a technical training course. They ask an AI assistant to solve 3(x - 4) + 5 = 20. The response is neat, correct, and easy to follow. Ten minutes later, the learner faces 4(x + 2) - 3 = 17 and cannot decide what to do first. Reading a solution created familiarity; it did not yet create independent performance.
A worked example can be an effective starting point for a novice, but the finished answer should not remain on screen for the entire session. The support needs to shrink. A useful sequence begins with one fully explained solution, removes selected steps from a second problem, and ends with a new problem that the learner solves without AI assistance.
This guide shows how to build that sequence with a general AI assistant while keeping a verified answer key outside the model. The aim is not to prove that one prompt teaches every learner. It is to create an observable handoff: each blank reveals whether the learner can choose and execute the next step.
Start with one skill and one trusted answer key
Choose a task narrow enough to observe. 'Teach me algebra' is not a session goal. 'Solve one-step and two-step linear equations while preserving equality' is. For the example session, the learner must distribute multiplication, combine terms, and apply the same operation to both sides. Factoring, quadratics, and word problems are outside the boundary.
Prepare three or four problems from a course text, instructor worksheet, or other source with checked solutions. Label one as the demonstration, two as completion problems, and one as the independent check. Do not ask the AI to invent the only answer key for material it will also grade. That makes one unverified system both author and judge.
The classic Sweller and Cooper algebra experiments compared conventional problem solving with instruction that made extensive use of worked examples. The result is relevant to beginners, not a command to show solutions forever. Later research on expertise reversal warns that guidance useful to novices can become redundant as knowledge grows.
- Target: one observable procedure or decision pattern.
- Boundary: concepts that belong in this session and concepts that do not.
- Key: solutions checked independently of the AI conversation.
- Exit task: a fresh problem that can be completed with no visible example.
Make the complete example explain decisions, not just arithmetic
A useful demonstration contains four layers: the current expression, the operation chosen, why that operation is legal or useful, and a quick check. For 3(x - 4) + 5 = 20, the explanation should say why distributing first exposes like terms, not merely jump from one line to the next.
Keep each line aligned with a single move. When an AI combines distribution, simplification, and subtraction in one leap, a learner may copy the line without seeing which rule produced it. Ask for the exact property beside the move, then compare the arithmetic and explanation with the prepared key.
Finish the demonstration by substituting the result into the original equation. This is not decorative reassurance. It models a verification habit the learner can later perform without the tutor. If the substitution fails, stop and repair the example before making variants from it.
Remove the decisions in a deliberate order
Fading is more precise than asking for an easier hint. In the first completion problem, show the setup and first transformation, then leave the simplification and final check blank. In the next problem, show only the original equation and a prompt such as 'Name the first operation and why.' The final problem contains no scaffold beyond the task itself.
The 2002 experiments by Renkl, Atkinson, Maier, and Staley tested a transition from complete examples through increasingly incomplete examples to independent problems. Their results supported near transfer in the studied settings, while far-transfer evidence was less clear. That boundary matters: solving a similar equation does not establish that the learner can model an unfamiliar word problem.
Remove one meaningful decision at a time. Deleting random arithmetic symbols produces a fill-in-the-blank puzzle, not necessarily practice in choosing a strategy. The blank should require the learner to supply a move, a reason, or a check that belongs to the learning target.
- Round 1: complete solution with a reason beside every move.
- Round 2: first move supplied; learner completes and checks the solution.
- Round 3: learner chooses the first move and explains it before continuing.
- Round 4: new problem, no visible example, followed by answer-key comparison.
Prompt for a sequence the learner must finish
Weak prompt: "Show me several examples of solving linear equations." The likely result is a page of complete solutions. It can be useful reference material, but it gives the learner no defined turn and no evidence of what they can do unaided.
Improved prompt: "I am a beginner practicing linear equations with distribution. Use the four instructor-checked problems below in the order provided. For Problem 1, explain one operation per line and name the property used. For Problem 2, show only the first operation, then stop for my work. For Problem 3, ask me to choose and justify the first operation before revealing any algebra. For Problem 4, provide no hints until I submit a full solution. Compare each submitted line with the supplied answer key, but do not copy the next keyed line into your feedback unless I ask for a hint. If my key and your calculation conflict, flag the conflict instead of silently choosing one."
Expected output at the start of Problem 2: the original equation, one verified first transformation, a single blank labeled 'Your next line,' and a request for the learner's reason. The assistant should then wait. A response that prints the remaining solution has broken the learning contract even if its mathematics is correct.
Put the problems and key in clearly separated blocks. Remove names, grades, unpublished test items, and other unnecessary private material before uploading anything. If a school or training provider restricts AI use, follow that rule and use approved practice material instead.
Use feedback that points backward before it points forward
When the learner submits a wrong line, first identify the earliest line that no longer follows from the previous one. Ask for a correction or offer a small cue tied to that move. Giving the entire correct solution converts the completion task back into reading.
A helpful first cue might say, 'Check how -4 changes when multiplied by 3.' A stronger second cue can ask the learner to write the distributive step with both products. Only after those attempts should the worked line appear. The hint ladder should be agreed in advance so that a fluent assistant does not over-help at the first hesitation.
AI feedback still requires scrutiny. In a 2024 PLOS ONE study across four mathematics subject areas, ChatGPT-generated hints were associated with learning gains in the tested design, but 32% of the initially generated hints failed the researchers' quality checks. The study used correctness filtering and a self-consistency method; a personal study session still needs a trusted key and human judgment.
Separate productive struggle from a broken example
A pause can mean the learner is choosing between plausible moves, which is useful work. It can also mean the example skipped a prerequisite, the blank is ambiguous, or the AI wrote an incorrect previous line. Do not interpret every delay as a need for more explanation.
Use a simple diagnosis. Ask the learner to state the goal of the next move. If they can state it but make a small arithmetic error, give a local cue. If they cannot state the goal, restore one earlier layer of support. If the supplied line conflicts with the answer key, stop the exercise and correct the material rather than coaching around the error.
A 2025 randomized study of a carefully engineered AI tutor in an undergraduate physics course reported stronger learning in its setting than the active-learning comparison. The system was not an unstructured chat: it used sequenced activities, detailed solutions, targeted prompts, and safeguards against inaccurate generation. The authors also cautioned that the result depended on that designed environment and should not be read as a case for replacing instructors.
- Learner gap: restore one prior step or prerequisite.
- Arithmetic slip: point to the exact operation without revealing later lines.
- Ambiguous blank: rewrite it so one intended decision is visible.
- Source conflict: pause, compare with the trusted key, and repair the sequence.
Promote the learner only after an unaided transfer check
The verification method is a closed-screen check. Give the learner a new equation with the same underlying moves but different surface details. Close the chat and hide every worked example. The learner writes each line, names the reason for at least one strategic move, substitutes the result into the original equation, and then compares the complete work with the independent key.
Score four observable items: first move chosen appropriately, transformations preserve equality, arithmetic is correct, and the final substitution works. A correct answer with an invalid middle step is not a pass. Neither is a solution that required reopening the example. Record which item failed and build the next completion problem around that decision.
Then test near transfer once: change the location of parentheses or place the variable on both sides if that variation belongs to the course target. Do not claim broad mastery from one successful item. Passing means the current scaffold can be reduced; it does not mean the learner has mastered every related problem.
The final page should belong to the learner
Worked examples are a bridge, not the destination. As the learner becomes faster and more accurate, long explanations and prefilled lines can become redundant. Shift the AI's role from demonstrator to checker, then from checker to an occasional source of new practice candidates that are reviewed before use.
Keep a small session record: the skill, source key, level of support used, first independent problem, errors found, and next variation. This makes progress visible without preserving the whole chat. It also reveals dependence: if every session remains at the fully worked stage, the support is not fading.
The useful question is not whether the AI produced a beautiful explanation. It is whether the learner can now make the next decision with less help, prove the work against a trusted key, and repeat the performance on a fresh problem. The best final artifact is not the model's answer. It is the learner's completed page.
Continue learning on JoyfulGrid
Frequently asked questions
How quickly should I remove steps from a worked example?
Remove support after the learner can explain and complete the current missing step accurately. If the same decision fails twice, restore one earlier layer, check the prerequisite, and try a new but comparable problem.
Can this method work for coding or spreadsheet formulas?
Yes, when the task has observable intermediate decisions and an independent way to test the result. Fade code lines, query clauses, or formula components deliberately, then run tests or compare outputs instead of trusting the AI's explanation alone.
Should the AI generate the practice problems?
It can draft candidates, but check that each problem is solvable, matches the target, and has a verified solution before using it. For assessed courses, prefer instructor-approved or textbook problems and follow the provider's AI policy.
What if the learner gets the right answer with different steps?
Compare the alternative path with the governing rules and key, not merely the final number. A different valid method can pass; an invalid transformation followed by a compensating error cannot.
Sources
- The Use of Worked Examples as a Substitute for Problem Solving in Learning AlgebraCognition and Instruction
Used for the foundational algebra experiments comparing conventional problem solving with extensive worked-example instruction.
- From Example Study to Problem Solving: Smooth Transitions Help LearningThe Journal of Experimental Education
Used for the complete-example to increasingly incomplete-example fading sequence and the reported near-transfer boundary.
- The Expertise Reversal EffectEducational Psychologist
Used for the caution that instructional guidance effective for inexperienced learners can lose value or become detrimental as expertise increases.
- ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skillsPLOS ONE
Used for the randomized study of AI-generated mathematics hints, its learning results, and the reported 32% initial hint quality-check failure rate.
- AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational settingScientific Reports
Used for the structured undergraduate physics tutoring study, its designed sequencing and accuracy controls, and the authors' implementation caveats.
