A learner opens a statistics chapter on correlation and causation. Before reading, they ask an AI assistant for a simple explanation. The answer mentions relationships, experiments, and confounding variables. Everything sounds reasonable. Ten minutes later, a new example asks why ice-cream sales and crime can rise together without one causing the other. The learner remembers the slogan "correlation is not causation" but cannot name the third variable or say what evidence would support a causal claim.

Now change the order. The learner first predicts an answer: perhaps ice cream causes people to become more aggressive. The guess is wrong, but it is concrete. When the source introduces warm weather as a plausible confounder, the correction has somewhere to land. The learner can compare two models rather than merely nod at a polished paragraph.

That is the useful role of AI pretesting. The assistant creates a small set of questions, protects the answer until the learner commits, and then helps compare the attempt with a trusted lesson. It does not turn every guess into learning, and the research does not support that broad claim. The workflow below uses the narrower evidence: prequestions can improve later memory for the material they target, while transfer and untested content need separate protection.

The useful moment happens before the answer

Pretesting means attempting a question before studying the answer. It differs from retrieval practice after study: there may be little relevant knowledge to retrieve, and errors are expected. The attempt can still activate related ideas, direct attention, and make the eventual correction more distinctive. In a 2009 series of six experiments, Nate Kornell, Matthew Hays, and Robert Bjork found that unsuccessful attempts improved later learning for fictional facts and weak word associations compared with reading the question and answer together.

The strongest recent summary is more precise than the slogan "guess first." A meta-analysis by Katherine St. Hilaire, Jason Chan, and Daeun Ahn reported a pooled specific effect of g = 0.54 across 97 studies for material that had been prequestioned. For untested material, the pooled general effect was g = 0.04 across 91 studies, with a confidence interval that included zero. In plain language, the attention benefit was much clearer for what the questions actually pointed toward than for the rest of the lesson.

This makes question selection an editorial decision. A weak pretest can teach the learner where to look while leaving the larger concept unattended. AI is helpful because it can produce variations quickly, but speed does not make its coverage complete or its answer key authoritative. The learner still needs a source boundary and a plan for content the prequestions did not sample.

Build three questions from one lesson boundary

Start with a defined section, not a subject name. For this session, the boundary is OpenStax Psychology 2e, section 2.3, especially its discussion of correlation, causation, confounding variables, experiments, and random assignment. Give the assistant the section title or a short excerpt. Do not ask it to roam across all of statistics.

Ask for three different prequestions. The first should test a central distinction. The second should require diagnosing a tempting but unsupported conclusion. The third should present a new surface example. Together they sample definition, reasoning error, and application. Keep the set short enough that the learner can answer each question in one or two sentences before seeing feedback.

Write coverage beside the questions. In this case the three rows might be relationship versus cause, possible third variable, and evidence from an experiment. After studying, scan the source once for major ideas that no row touched. That final coverage scan matters because the meta-analysis found little evidence that a prequestion automatically improves learning of everything nearby.

  • One bounded source: a chapter section, lecture segment, or instructor handout.
  • Three question roles: distinguish, diagnose, and apply.
  • A written attempt: answer plus confidence and one reason.
  • A coverage note: concepts tested, concepts not tested, and concepts outside scope.

Tell the assistant to protect the first attempt

Weak prompt: "Teach me correlation and causation, then quiz me." The explanation arrives first, so the quiz measures very recent exposure. The assistant also chooses the lesson boundary, the important claims, and the grading standard without showing where any of them came from.

Improved prompt: "Use only the source excerpt I provide. Create three short prequestions: one central distinction, one diagnosis of a flawed conclusion, and one new scenario. Do not show answers, hints, key terms, or a summary yet. Ask one question at a time and wait for my answer. Record my answer and confidence. After all three attempts, return a table with: my claim, supported / contradicted / not covered, the exact source paragraph that resolves it, and one question I should answer after studying. Do not treat your own background knowledge as source evidence."

Expected output: the assistant first asks something like, "A city finds that umbrella sales and traffic collisions rise on the same days. What can the relationship establish, and what third factor would you examine before making a causal claim?" It stops there. After the learner answers, it records the wording without correcting it. Only after the three attempts does it point to the supplied paragraph on confounding and ask the learner to revise the causal claim.

The stop instruction is essential. A hint such as "think about weather" gives away the useful comparison before the learner has generated a model. If the assistant leaks an answer, discard that item as a prequestion and replace it with a fresh scenario. Do not pretend the attempt was blind when it was not.

A twenty-four-minute correlation session

Minutes 0-4: paste or identify the bounded OpenStax section and request the three questions. On paper, create four columns: question, first answer, confidence from 0 to 100, and reason. Confidence is not a score. It helps distinguish an uncertain guess from a misconception held with conviction.

Minutes 5-9: answer each question without searching, opening the textbook, or asking for a hint. Use complete claims, not single words. A first answer might say, "The relationship proves umbrella purchases distract drivers. Confidence: 35. Reason: both variables rise together." The low confidence does not make the causal leap harmless; it simply records the learner's state before correction.

Minutes 10-16: read the source section with the three questions visible. Mark the sentence that confirms, narrows, or contradicts each first answer. For the umbrella case, the source principle is that a correlation describes a relationship but does not by itself establish cause and effect; rain is a plausible common factor. The invented example comes from the study exercise, while the rule comes from the source.

Minutes 17-20: let the assistant build its comparison table, then inspect every locator yourself. Rewrite each answer in your own words and add one condition. A stronger revision might say, "The data establish an association. Rain could increase both umbrella purchases and difficult driving conditions, so the correlation alone does not show that buying an umbrella causes collisions."

Minutes 21-24: close the source and the chat. Answer a fourth scenario that was not in the pretest. Then write one question the evidence still cannot answer. This exit task converts a targeted attention exercise into an independent check rather than ending with the assistant's correction on screen.

Keep feedback close, narrow, and source-linked

An error should not sit uncorrected through a long conversation. The workflow delays feedback only until the small pretest is complete, then places the first answer beside the relevant source passage. That comparison is the learning event. A generic message such as "not quite" or a fresh lecture makes the original reasoning harder to inspect.

Ask the assistant to label each claim, not the whole response. A learner may correctly identify a correlation and still invent the direction of causation. Another may name a plausible confounder but claim it has been proven by the data. Claim-level labels preserve the difference. They also fit naturally with an AI error log: record the exact decision that failed, the evidence that corrected it, and the new case that tested the repair.

Treat the assistant's table as navigation, not grading authority. Open the cited paragraph and confirm that it says what the label claims. If the source is silent, mark the row not covered. Fluent background knowledge from the model does not become part of the assigned lesson merely because it is relevant.

Notice the attention tunnel and the answer leak

The attention tunnel appears when the learner studies only the sentences that answer the three prequestions. Research on prequestions has repeatedly raised this tradeoff, and the recent meta-analysis found the general benefit for untested material was close to zero. Counter it with a coverage scan: after resolving the questions, list the section headings and write one sentence about each untested idea.

The answer leak appears when the question contains the key term, the hint supplies the mechanism, or a multiple-choice distractor makes the intended distinction obvious. The remedy is procedural: label the item exposed and replace it. The confident-wrong-key failure is different. The assistant may produce a plausible correction that the source does not support. Require paragraph locators, then perform a direct source check before saving the answer.

The trivia illusion happens when a successful workflow for discrete facts is stretched into a claim about deep conceptual learning. Hausman and Rhodes found that conceptual pretests did not enhance concept learning from reading in their experiments; feedback improved repeated conceptual items, but the authors warned that this could reflect memorization rather than conceptual understanding. The safer design includes a new scenario and asks for reasoning, not just the previous answer.

The pretest-only failure turns an opening activity into the entire study method. In one open-access experiment, Alice Latimier and colleagues found both pretesting and post-testing benefits after seven days, but post-testing produced better retention and transfer to untrained questions. Pretesting can focus the first read. It does not replace retrieval after study, spaced review, or later application.

Verify twice: once against the source, once after a delay

The first verification pass is documentary. For every corrected claim, save the source title, section, and the sentence or paragraph that supports it. Separate direct statements from inferences. If the assistant supplies an additional concept such as selection bias, causal graphs, or regression controls, place it in a needs-another-source row rather than quietly expanding the lesson.

The second pass is behavioral. One or two days later, answer two fresh cases without the old chat: one similar case and one with a misleading surface detail. For example, distinguish an association between exercise and mood from a randomized experiment that assigns an activity, then identify what random assignment helps control and what the reported result still does not prove beyond the study conditions.

Score the reasoning in parts: relationship described accurately, causal language justified or withheld, alternative explanation named, design evidence identified, and uncertainty preserved. If a part fails, reopen only the relevant source section. This is stronger evidence of usable learning than remembering the wording of the original umbrella example.

  • Source pass: can every correction be traced to the assigned material?
  • Near case: can the learner apply the distinction with new nouns and numbers?
  • Changed-design case: can the learner notice when the evidence type changes?
  • Delayed pass: can the reasoning be rebuilt without the AI transcript?

Finish when the question has done its job

A prequestion is temporary scaffolding. It has succeeded when it makes a knowledge gap visible, guides attention to a relevant passage, and leaves behind a corrected rule that survives a new case. Keeping the whole chat or generating twenty more questions does not improve that evidence.

Use AI for the parts it can make easier to repeat: creating short blind scenarios, recording first answers, aligning claims with a supplied excerpt, and varying the delayed check. Keep the important judgments outside the model: whether the source is trustworthy, whether the question coverage is balanced, whether the locator is accurate, and whether the learner can reason alone.

The order is simple but demanding: attempt, study, compare, verify, and return later. The learner gets the first word and the final test. The assistant works in the middle, where a wrong but visible idea can become a precise correction instead of disappearing beneath a fluent explanation.

Continue learning on JoyfulGrid

Frequently asked questions

Will guessing first make me remember the wrong answer?

An unsupported guess should be corrected promptly and explicitly against the learning source. The cited evidence shows benefits under some conditions, not a guarantee for every task; keeping the first answer beside the correction and then testing a new case reduces the chance that an unexamined guess becomes the study note.

How many prequestions should I use?

For one short section, three is a practical starting point: a central distinction, a diagnosis, and an application. Add a coverage scan rather than generating a long quiz that narrows attention to dozens of model-selected details.

Is pretesting better than a quiz after studying?

Do not treat them as substitutes. Pretesting can focus attention before study, while post-study retrieval checks what can be produced after learning. One direct comparison found stronger retention and transfer from post-testing, so this workflow deliberately includes both an opening attempt and a delayed test.

Can I use this with material I know nothing about?

Yes, if the questions are low-stakes, bounded, and followed by reliable corrective study. Record uncertainty, avoid sensitive or safety-critical guessing, and judge success by the verified revision and later application rather than by the accuracy of the first attempt.

Sources

  1. Guessing as a Learning Intervention: A Meta-Analytic Review of the Prequestion EffectPsychonomic Bulletin & Review

    Used for pooled specific and general prequestion effects, the tested-versus-untested distinction, and the need to avoid assuming broad lesson-wide benefits.

  2. Prequestioning and Pretesting Effects: a Review of Empirical Research, Theoretical Perspectives, and Implications for Educational PracticeEducational Psychology Review

    Used for definitions, classroom evidence, question-format considerations, and the limits of current evidence on transfer and general benefits.

  3. Unsuccessful Retrieval Attempts Enhance Subsequent LearningJournal of Experimental Psychology: Learning, Memory, and Cognition via PubMed

    Used for the original six-experiment evidence that an unsuccessful attempt can improve later learning for the factual and associative materials studied.

  4. When Pretesting Fails to Enhance Learning Concepts from Reading TextsJournal of Experimental Psychology: Applied via PubMed

    Used for the conceptual-learning boundary and the warning that gains on repeated items may reflect memorization rather than broader understanding.

  5. Does Pre-Testing Promote Better Retention Than Post-Testing?npj Science of Learning

    Used for the direct seven-day comparison of pretesting, post-testing, and rereading, including the stronger retention and transfer results for post-testing in that experiment.

  6. 2.3 Analyzing FindingsOpenStax Psychology 2e

    Used as the trusted content boundary for the correlation, causation, confounding-variable, experiment, and random-assignment learning scenario.