The Testing Effect: Why Quizzes Beat Rereading [2026]
The testing effect proves quizzing builds stronger memory than rereading. See how quiz duels and spaced repetition apply it.

You read a chapter twice. Your friend took one quiz on it instead. A week later, your friend remembers about twice as much as you do.
That gap has a name: the testing effect. Pulling a fact out of memory strengthens it far more than looking at it again. In the experiment that anchors the modern research, Karpicke and Roediger (2008) checked on learners a full week after study, and the group that kept testing itself still held 80% of the material while the group that kept restudying without tests was down to 36%.
I run a quiz app, so you should know I have a commercial interest in this finding. But the finding came long before the app: LearnClash’s duels and Solo rounds are built as forced retrieval because a century of research says retrieval is the part that sticks. This article covers that research, where the effect breaks down, and what we did with it. Test it yourself with a quiz duel on any topic →
| Testing Effect in LearnClash | |
|---|---|
| The core move | Every question is a recall attempt |
| Duels | 18 questions per match, three per round across 6 topics |
| Solo mode | 6-question rounds on any topic you pick |
| Spaced repetition | Missed questions return in 7 days; the second checkpoint sits at 90 |
| Difficulty | Easy, medium, hard, scaled to your rating |
What Is the Testing Effect?
The testing effect is the finding that retrieving information from memory builds stronger long-term retention than passively reviewing the same material. Reaching for an answer, even reaching and coming up empty, lays down more durable pathways than reading that answer five times over. The reach is the work.
Rereading builds a single fragile trace. Recall builds a branching web of links.
Rereading fails precisely because it feels so productive. The text looks familiar, the ideas seem clear, and you close the book confident. Psychologists call this the fluency illusion: the ease of reading tricks your brain into filing the material as stored when it is not.
Knowing and recalling are different tasks. You can know a face and still blank on the name attached to it, and an exam only ever asks for the name.
“Testing is a powerful means of improving learning, not just assessing it.” Roediger & Karpicke, Perspectives on Psychological Science (2006)
So the practical split looks like this. Reread your notes three times and you have practiced recognition. Close the notes and write down what you can actually recall, and you have run a memory test, the thing that builds the trace.
Why Does Quizzing Yourself Beat Rereading?
Quizzing beats rereading because retrieval forces your brain to rebuild the route to the answer, and rebuilt routes are the ones that survive. Rereading skips the rebuilding entirely. It leaves you with the comfortable sense that the material is familiar while the path you would need under exam pressure never gets laid.
Rereading looks better minutes after study, then the lines cross within days. One-week figures from Karpicke & Roediger (2008); the early crossover pattern appears in their 2006 experiments.
The landmark proof came from a 2008 study in Science. Karpicke and Roediger had students learn Swahili-English word pairs, and once a student got a pair right, it was dropped from further study sessions, from further tests, from both, or from neither. Everyone took a final exam one week later.
| Condition | Recall after 1 week |
|---|---|
| Kept studying the pairs, stopped testing on them | ~36% |
| Stopped studying the pairs, kept testing on them | ~80% |
Whether students kept studying made no measurable difference. Only the testing mattered.
Two large reviews since have put numbers on how reliable that result is:
- Rowland (2014) pooled 61 testing-effect studies and landed on an effect size of g = 0.50 overall, rising to g = 0.73 when feedback followed each attempt.
- Adesope and colleagues (2017) ran the math over 272 effect sizes from 188 studies and found g = 0.61, with the advantage widening over time to g = 0.82 at delays of one to six days.
That last figure deserves a second look. The longer the gap between practice and the real test, the wider testing’s margin over rereading. Cramming assumes the exact opposite.
A Century of the Same Result
Arthur Gates ran the first formal study back in 1917, and reciting from memory beat rereading in his classroom experiments. The groundwork is older still: Hermann Ebbinghaus mapped the forgetting curve in 1885 and showed how quickly new material drains away without review. The 2000s wave from Roediger, Karpicke, and Bjork consolidated a scattered century of observations into one of the firmest findings in learning science.
Over 100 years of research. The verdict never changed: testing beats rereading.
| Year | Researcher | What They Found |
|---|---|---|
| 1885 | Hermann Ebbinghaus | Mapped the forgetting curve using nonsense words |
| 1917 | Arthur Gates | First study: reciting beats rereading |
| 1939 | Herbert Spitzer | Proved it in real classrooms (3,600 students) |
| 2006 | Roediger & Karpicke | Modern landmark: a week out, tested students beat rereaders |
| 2011 | Karpicke & Blunt | Testing beats concept mapping, even on linking questions |
| 2013 | Dunlosky et al. | Rated practice testing “high utility” (only 2 of 10 methods scored that) |
| 2021 | Agarwal et al. | Classroom review: testing works across ages and subjects |
One detail in this history still trips people up: students predict the wrong winner. In Roediger and Karpicke’s own experiments, the students who restudied rated their future recall highest, and a week later the tested groups beat them. The fluency illusion runs that deep.
Robert Bjork gave the pattern a name in 1994: desirable difficulties. Methods that slow early learning, such as testing, spacing, and mixing topics, tend to produce the strongest recall later, or in the shorthand Bjork and Bjork have used since, conditions that slow apparent learning often optimize long-term retention and transfer. The testing effect is the most studied member of that family.
How the Effect Works in Your Brain
The mechanism has a name: retrieval-induced strengthening. Attempting recall reactivates the relevant memory trace and lays down additional routes toward the same fact, so the next attempt has more paths available than the one before.
Recall forces your brain to rebuild the path to the answer. Rereading just shows you the path exists.
The Recall Process
When a question lands, your brain hunts through stored links for the answer, and that search strengthens every path it touches. Karpicke and Smith (2012) traced the benefit to retrieval-specific mechanisms rather than extra exposure time. The search itself does the building, not the seconds spent looking at the page.
Even failed recall helps. Butterfield and Metcalfe (2001) documented the hypercorrection effect: answer with high confidence, get it wrong, and you become more likely to remember the correction than if you had guessed timidly. The jolt of being wrong while feeling sure flags the fix as high priority. Our misconception index catalogs that feeling at scale: the 20 wrong answers LearnClash players were most confident about, seven of which outpolled the correct answer.
Why Harder Recall Creates Better Memory
Not every retrieval is worth the same. When an answer refuses to surface right away and you have to genuinely grind for it, the memory gain measurably exceeds what an effortless fact delivers. That traces straight back to Bjork’s desirable-difficulty work.
There is a ceiling, though. Recall attempts on material you never learned in the first place produce frustration, not strengthening. The productive zone sits where retrieval takes real effort and still succeeds most of the time.
How We Built LearnClash Around It
Every design review for LearnClash eventually hits the same question: does this make the player retrieve, or does it let them recognize? Duels put 18 recall attempts in front of you per match, three per round across 6 different topics, half of them picked by your opponent. Solo mode runs 6-question rounds on any topic you pick. Both feed the same spaced repetition schedule.
Two modes, same science. Duels add ranked pressure. Solo adds precise timing.
Duels: Recall With Something Riding on It
A duel spreads its questions across 6 different topics. You and your opponent alternate picking them, each choosing from a small hand of options the app deals, and no topic repeats within a match. Half the rounds run on your opponent’s picks, topics you never chose, which adds topic mixing, one more desirable difficulty. You answer under a 45-second clock against a real opponent matched near your skill level. The rating is computed with Glicko-2, and new players start with a high rating deviation, so early duels swing hard and every answer counts.
The stakes are not decoration. Focus is the fuel retrieval runs on, and a ranked answer gets a level of attention an unscored flashcard rarely does. As your rating climbs, the questions scale up with you, which keeps recall effortful without tipping into hopeless.
No Points for Speed
One scoring choice runs against every trivia-app instinct: answering fast earns you nothing in LearnClash. Every question, duel or Solo, runs on the same 45-second clock, and a correct answer in second 44 scores exactly what it scores in second 2. We record response times for analytics, but they never touch the score.
I made that call because speed bonuses pay out for recognition, the shallow process this whole article is about escaping. The timer exists to force commitment before the reveal, not to reward reflexes. If you need the extra thirty seconds to drag an answer up from somewhere deep, that grind is the desirable difficulty doing its work, and the game should not tax you for it.
What Happens After You Miss
Wrong answers sting, and the sting is doing work. Butterfield and Metcalfe’s hypercorrection research says high-confidence errors get corrected at the highest rates, so LearnClash reveals the correct answer the moment you commit, when the surprise peaks. Then the spaced repetition system takes over.
| Stage | What it means | Next scheduled review |
|---|---|---|
| Learning | New question, or missed on a recent attempt | 7 days |
| Known | Answered correctly once, off cooldown | 90 days |
| Mastered | Cleared the 90-day checkpoint too | None; the question retires from rotation |
A miss drops a question one stage rather than resetting it to zero, and two spaced correct answers retire a question climbing from Learning; a Mastered question that slips re-enters at Known and needs only its fresh 90-day check to retire again. While fact-checking this update against our own code I caught the earlier version of this article describing that schedule wrong: it promised Mastered questions a 90-day review, and they get none. Retiring what you have proven twice is the point, because review time should stay reserved for the facts still fighting you. The stage ladder itself descends from the Leitner system, the 1972 cardboard-box schedule where correct cards climb toward rarer review and missed cards fall back.
Testing plus spacing beats either one alone. Dunlosky et al. (2013) reviewed ten common study methods and handed out exactly two “high utility” ratings, one to practice testing and one to spaced practice. Pairing them was the design goal from the start.
Where the Effect Breaks Down
The testing effect is robust, not magic. Material so new there is nothing to retrieve, a learner with no reason to try, recall with no feedback afterward: each of these drains the benefit in the lab.
The testing effect isn’t magic. It has limits. LearnClash is designed around them.
| Failure Case | Why It Fails | LearnClash Fix |
|---|---|---|
| Facts too new | No recall attempt possible | Three levels: easy, medium, hard |
| Low drive or focus | Shallow work, no real effort | Ranked stakes on every answer |
| No feedback | Errors go unfixed, may lock in wrong answers | Instant answer reveal after each question |
| Questions too hard | Frustration blocks useful struggle | Matching prefers rivals at a close skill level |
A 2026 study in Frontiers in Psychology makes the motivation point uncomfortably well. The authors ran a standard retrieval-practice design on the crowdworking platform Prolific, with attention checks, feedback, and fair pay, and still found no testing effect at the delayed test. Their reading is not that the effect is fragile, but that crowdsourced settings struggle to produce the sustained engagement it depends on.
That study is the strongest argument I know for making recall practice feel like play rather than homework. Ranked stakes are our answer to the engagement problem: when a rating is on the line, each answer gets real effort, and real effort is the ingredient those null results were missing.
Test the testing effect yourself with a history duelTesting Versus the Other Study Methods
Dunlosky et al. (2013) graded ten common learning techniques, and practice testing finished in the top tier while rereading and highlighting landed at the bottom. That review remains the standard reference for study-method advice.
Six study methods ranked by long-term recall. Only two earned “high utility” from Dunlosky’s review.
| Study Method | Long-Term Recall | Dunlosky Rating |
|---|---|---|
| Highlighting | Very low | Low |
| Rereading | Low (36% at 1 week) | Low |
| Summarizing | Medium | Low |
| Concept mapping | Medium | Not rated |
| Practice testing | High (80% at 1 week) | High |
| Testing + spacing | Strongest combination | Both high |
Karpicke and Blunt (2011) found an even more striking result in Science. They pitted recall practice against concept mapping, a method widely seen as better for deep understanding, and testing won on factual recall and on questions that required linking ideas across the text. Even when the final task was to draw a concept map, students who had practiced recall beat students who had practiced mapping.
Students surveyed beforehand backed concept mapping. So recall practice is not a shallow memorization trick; it builds the kind of flexible, connected knowledge the mapping crowd assumed they owned.
There is also an honesty angle. Rereading lets you keep flattering illusions about what you know, and a test result does not negotiate. That is the same reason a rating gives you a cleaner read on what stuck than quiz apps built purely for fun can offer.
Five Rules, No App Required
The whole effect compresses into one swap: trade passive review for active self-quizzing. Five rules cover it.
- Close your notes and recall first. Write down everything you can before rereading anything. The gaps you find are the to-do list.
- Quiz yourself before you feel ready. Struggling for an answer before mastery is the desirable difficulty that builds the trace.
- Mix your topics. Alternating subjects forces your brain to tell similar facts apart instead of coasting on context.
- Space your sessions across days. The forgetting curve is steep, and spacing is the only proven counter.
- Add stakes. Recall without focus produces weak effects, which is exactly what the Prolific null results show. A study partner, a bet, or skill-matched competition all work.
Five steps from passive to active to ranked. Each step boosts the testing effect.
If you would rather have all five run automatically, that is what the app is for: duels supply the stakes and the topic mixing, Solo mode supplies the self-quizzing, and the SRS handles spacing without a spreadsheet. If you have ever quizzed yourself from a list like our science trivia questions before an exam, you have already felt the effect. LearnClash’s real addition is the scheduling and the rival.
For the wider toolkit around retrieval, see how to study effectively, the deadline tactics in how to memorize fast, or how the two biggest study apps score on retention in our Kahoot vs Quizlet comparison.
Explore more learning science articlesFrequently Asked Questions
Does quizzing yourself actually work better than rereading?
Yes. Karpicke and Roediger's 2008 study in Science showed students who continued testing recalled 80% of material after one week, while those who stopped testing recalled only 36%. LearnClash applies this in every quiz duel and practice session, turning each answer into a memory-strengthening retrieval event.
What is a real-world example of the testing effect?
Answering a quiz question incorrectly, seeing the correct answer, and then remembering it weeks later. LearnClash creates this in every duel: you attempt retrieval under time pressure, see the answer immediately, and the spaced repetition system schedules follow-up reviews at 7-day and 90-day intervals.
Is the testing effect the same as retrieval practice?
The testing effect is the observed result. Retrieval practice is the technique that produces it. Forcing yourself to recall information rather than passively review it strengthens the memory trace. LearnClash quiz duels and practice mode are both forms of retrieval practice that trigger the testing effect.
How many times do you need to quiz yourself to remember something?
The research consistently favors a few successful retrievals spread over expanding intervals rather than many crammed together. LearnClash's 3-stage spaced repetition system settles on two spaced checkpoints: a new or missed question returns after 7 days, a correct answer sets a 90-day checkpoint, and clearing that one marks the question Mastered and retires it from the review rotation. A miss drops a question back one stage rather than resetting it to zero.
Can the testing effect work for all subjects?
Yes. Studies confirm the testing effect across languages, sciences, history, medical training, and general knowledge. LearnClash lets you quiz yourself on any topic at three difficulty levels, applying retrieval practice to whatever you want to learn.
