A biology learner reads that very high temperature can reduce enzyme activity. The sentence seems familiar, so they ask an AI assistant to explain it. The response is fluent: heat changes the enzyme, the active site no longer fits the substrate, and the reaction slows. After reading twice, the learner feels ready.
A blank-page question exposes the gap: why can a modest temperature increase speed a reaction while a larger increase can slow it? The learner remembers the words "more collisions" and "denaturation," but cannot connect temperature, molecular motion, protein structure, binding, and reaction rate. The explanation was available to read, yet the mechanism was never assembled by the learner.
Self-explanation reverses the order. The learner first states how and why each step follows from the last. AI is then useful as a questioner, a contrast generator, and a checker against supplied material—not as the voice that performs the explanation before the learner has tried. This workflow keeps that sequence visible from the first sentence to the final transfer question.
A polished explanation can hide borrowed understanding
Self-explanation is more specific than summarizing. A summary compresses what a source says. A self-explanation generates an inference: it connects a step to a principle, supplies a missing cause, reconciles a contradiction, or states the condition under which a claim holds. In the 1989 study that established the self-explanation effect, Michelene Chi and colleagues analyzed how students studied worked mechanics examples. Stronger learners generated more explanations that related solution steps to principles and later depended less on the original examples.
The broader evidence is encouraging but should not be turned into a universal promise. A 2018 meta-analysis by Kiran Bisra and colleagues combined 69 effect sizes from 64 research reports and estimated an overall weighted effect of g = .55 across varied tasks and settings. A mathematics-focused meta-analysis by Bethany Rittle-Johnson and colleagues found small-to-moderate benefits for immediate procedural, conceptual, and transfer outcomes, while evidence for classroom use and delayed retention was more limited. It also found stronger immediate effects when learners received scaffolding for high-quality explanations.
That nuance determines the role of AI. A chatbot can supply targeted prompts whenever a causal link is missing, but reading its complete explanation is not the same activity as generating one. The useful unit is not a long answer. It is one learner-written link followed by one question that makes the link more precise.
Start with a claim strip, not an empty chat
Choose a small piece of trusted material and extract four to six claims that the learner is allowed to use. For the enzyme lesson, OpenStax explains that enzymes lower activation energy, substrates bind at an active site, the active site's chemical environment supports specific interactions, and temperature or pH outside a suitable range can impair binding or denature the enzyme. Those statements become the lesson boundary.
Next, write one question that requires a mechanism rather than a definition: "Why can warming increase an enzyme-catalyzed reaction rate at first, yet excessive heat reduce it?" Under it, draw three boxes labeled condition, mechanism, and observable result. The first attempt must be completed without AI. An honest blank box is more useful than a borrowed sentence because it shows where the next prompt should aim.
Keep confidence beside each link, not beside the whole topic. A learner may be confident that temperature affects molecular motion but unsure how structural change affects substrate binding. Marking those separately stops a general feeling of familiarity from standing in for a tested causal chain.
- Claim strip: four to six source-backed statements, with page or section markers.
- Mechanism question: asks how one condition produces one result.
- First chain: condition → intermediate cause → observable result, written without assistance.
- Confidence marks: sure, uncertain, or missing for each arrow.
Make the assistant interrogate the arrows
Weak prompt: "Explain how temperature affects enzymes in simple terms." This asks the assistant to select the scope, build the mechanism, choose the vocabulary, and judge what counts as simple. The learner can agree with the result without revealing whether any link was understood.
Improved prompt: "I will paste a short source excerpt and my own three-link explanation. Do not rewrite or finish my explanation. Inspect one arrow at a time. For the first unsupported, vague, or contradictory arrow, ask one question that makes me name the mechanism. Wait for my reply. Then label my revised link as supported, partly supported, contradicted, or not covered by the excerpt, and quote no more than one short phrase from the excerpt as evidence. Continue only when I say next. End by asking me to predict a changed case."
Expected output: if the learner writes "higher temperature → the enzyme works faster," the assistant should not replace it with a lecture. It might ask, "What changes at the molecular level that could increase successful enzyme-substrate encounters before the protein's structure is disrupted?" After the learner answers, the assistant should distinguish what the supplied source supports from what needs another reference. The response is deliberately incomplete because the learner owns the next sentence.
The one-question limit matters. A list of six hints can function like an answer key, especially when the later questions reveal the intended sequence. Ask the model to stop after one diagnostic question and require the learner to revise the arrow in their own words before moving on.
Separate source evidence from the model's teaching choices
The source and the assistant have different jobs. The source establishes the domain claims. The assistant chooses which question to ask about the learner's current wording. If it adds a term such as collision frequency, kinetic energy, conformational change, or reversible inhibition, label that term as an addition until it is checked against the course material or another suitable source.
This separation is especially important because explanation quality is not just factual accuracy. The 2025 ELI-Why study evaluated language-model explanations for learners at elementary, high-school, and graduate levels. In its human studies, GPT-4 explanations matched the intended educational background 50 percent of the time versus 79 percent for lay human-curated explanations, and users rated the model explanations as less suited to their informational needs on average. Those figures describe the models, benchmark, and raters in that study; they are not a universal score for every AI explanation. They do show why "make it suitable for me" is not a sufficient control.
Give the assistant evidence-handling rules instead: do not claim support when the excerpt is silent; preserve conditions and qualifiers; distinguish a direct statement from an inference; and point to the exact sentence or section that supports a label. Then inspect the cited passage yourself. The model can route attention, but it cannot turn an uncited addition into course evidence by sounding certain.
A twenty-minute enzyme session
Minutes 0–4: read the selected OpenStax section and build the claim strip. Write the mechanism question, then make a first chain. A plausible attempt might be: "moderate warming → particles move more → substrate meets the active site more often → reaction rate rises; excessive heat → bonds supporting protein shape are disrupted → active-site interactions become less suitable → activity falls." Mark any link whose wording goes beyond the excerpt.
Minutes 5–11: paste the excerpt and the chain into the improved prompt. Answer only one diagnostic question at a time. When the assistant flags a link as partly supported, decide whether to narrow it, locate another trusted source, or leave it unresolved. Do not let the conversation expand into every factor that affects enzymes; the session question controls the boundary.
Minutes 12–15: hide the chat and rewrite the chain as a paragraph from memory. Underline each causal connector—because, therefore, which changes, so that—and ask whether the sentence on each side genuinely supports that connector. Compare the paragraph with the claim strip and correct only with evidence.
Minutes 16–20: answer a changed case. For example: "Two samples use the same enzyme and substrate. Sample A rises from 20°C to 30°C; sample B is held at a much higher temperature and then cooled. What different mechanisms could explain their reaction rates, and what observation would help distinguish them?" The goal is not to guess a universal temperature curve. It is to state conditional predictions, identify missing information, and apply the mechanism without copying the original wording.
- Product: one claim strip, one learner-written chain, and one rewritten paragraph.
- AI contribution: diagnostic questions and evidence labels, not the finished mechanism.
- Independent evidence: the supplied textbook section or another instructor-approved source.
- Exit evidence: a prediction for a changed condition with assumptions made explicit.
Notice when the conversation has taken over
The completion leak happens when a supposedly Socratic question contains the missing answer: "Could heat have broken the bonds that maintain the active site's shape?" The learner only has to agree. Ask for a question that identifies the type of missing link without naming its content, or have the assistant offer two competing questions rather than a proposed fact.
The paraphrase loop happens when the learner replaces one textbook sentence with synonyms. A causal connector is present, but no new inference has been made. Require each arrow to answer one of three tests: what changed, why did it change, or how would the result differ if this link were absent? If none applies, the sentence may be description rather than explanation.
The invented-mechanism failure occurs when the assistant supplies a plausible biochemical detail not present in the lesson. Do not automatically delete it, but do not absorb it into notes either. Put it in a parking column labeled "needs a source." The confidence-mirroring failure occurs when the assistant accepts polished learner language without challenging a missing condition. Ask it to inspect every causal verb and identify the first leap, even when the prose sounds fluent.
The endless-dialogue failure is subtler. The learner keeps answering prompts but never reconstructs the mechanism alone. Set a hard handoff after three to five questions. Close the chat, write from memory, and use the independent source to judge the result. If the chain collapses, reopen only at the first missing arrow.
Keep a compact explanation record
Save one page rather than the transcript. At the top, record the mechanism question and source boundary. In the middle, keep the final chain with one evidence marker beside every arrow. At the bottom, record the changed case, your prediction, the check you used, and the first link that failed if the prediction was wrong.
A useful evidence marker can be simple: D for directly stated, I for a reasonable inference from stated claims, X for contradicted, and U for unresolved. The letters are not grades. They prevent an inference from slowly turning into a remembered quotation and make the next review session selective.
Return a day or two later and reconstruct the chain without the saved page. Then change one variable or representation. Draw the mechanism, explain it aloud, or diagnose a deliberately flawed explanation. Delayed retention is less studied than immediate self-explanation effects, so the later check is not optional proof of a guaranteed benefit. It is your direct evidence that this particular explanation remained usable.
Finish when the mechanism can travel
The purpose of self-explanation is not to produce the longest account. It is to expose the links that control a prediction. A finished explanation states the relevant condition, connects each step with a supported mechanism, preserves important limits, and produces a defensible answer when the surface details change.
AI can make that practice easier to run: it can notice a vague arrow, ask for a missing cause, generate a contrasting case, and keep the interaction moving. Its best contribution is often a well-timed absence—the refusal to complete the thought before the learner has generated it.
When you can rebuild the chain without the chat, trace its claims to the source, and use it on a changed example, stop. The explanation no longer belongs to the assistant's response. It has become a model you can inspect, revise, and use.
Continue learning on JoyfulGrid
Frequently asked questions
Is self-explanation just the Feynman technique?
They overlap, but self-explanation research usually focuses on generating inferences that connect new information, principles, and problem steps. Explaining simply to an imagined audience can help, but simplicity alone does not prove that the causal links are supported.
Should I speak or write my explanation?
Either can reveal gaps. Writing makes each arrow and evidence marker easier to inspect; speaking can reduce friction and expose where you hesitate. For important material, record a short spoken attempt and then write the final causal chain.
Can the AI grade my final explanation?
Use it for a provisional critique, not as the only judge. Require evidence labels tied to the material you supplied, then verify the claims and conditions against the original source, an instructor key, or another qualified reference.
Sources
- Self-Explanations: How Students Study and Use Examples in Learning to Solve ProblemsCognitive Science
Used for the original analysis of learner-generated explanations while studying worked mechanics examples and later example-independent problem solving.
- Inducing Self-Explanation: a Meta-AnalysisEducational Psychology Review
Used for the definition of self-explanation, the 69 included effect sizes from 64 reports, the weighted mean effect, and the range of instructional conditions studied.
- Promoting Self-Explanation to Improve Mathematics Learning: A Meta-Analysis and Instructional Design PrinciplesZDM – Mathematics Education
Used for evidence on immediate mathematics outcomes, the value of scaffolding high-quality explanations, and cautions about classroom and delayed-retention evidence.
- ELI-Why: Evaluating the Pedagogical Utility of Language Model ExplanationsAssociation for Computational Linguistics
Used for the 2025 benchmark and human-study findings about how well generated explanations matched intended educational levels and learner needs.
- 6.5 Enzymes – Biology 2eOpenStax
Used as the trusted source boundary for the enzyme learning scenario, including active sites, activation energy, temperature, pH, binding, and denaturation.
