Rethinking Retrieval Practice: Remembering Is Not Knowing
Retrieval practice works best as the maintenance of meaningful learning, not its creation.
I’ve been talking recently about how retrieval practice is going wrong in a lot of classrooms. A lot of practice which invokes the science of learning is built on a simple conviction: if pupils remember more, then they will do more. In other words, strengthen the memory trace and performance will rise. However, two recent studies on retrieval practice put a crack in that picture and and echo older warnings from Ryle and others that knowing a rule and being able to use it are different kinds of achievement. In both of these new studies, pupils and students remembered more but yet did no better on application tasks.
Knowing That, Knowing How
In The Concept of Mind, published in 1949, Gilbert Ryle drew a distinction between two kinds of knowing. When we teach a child a rule, we are doing something in the domain of knowing-that. When we want her to be able to use the rule, we are asking for something in the domain of knowing-how. The two are related; but they are not the same; and nothing in the logic of the first guarantees the delivery of the second.
A child who can recite the rules of chess is not thereby a chess player; a child who can describe the strokes of swimming is not thereby a swimmer. Knowing that and knowing how are separable epistemic achievements. One can travel a long way in the first without taking a single step into the second. The distinction is philosophical in origin but has real implications for instructional design.
Some recent studies I read reminded me of this and also really underlined a lot of why retrieval practice is going wrong in many classrooms. A new study/doctoral thesis on retrieval practice reminded me of Ryle’s classic work and puts this old distinction through an empirical test. A group of Dutch fourth‑graders, spread across seven regular primary classes, spent a week learning to spell their language’s trickiest verbs. All pupils received a scripted lesson on the relevant rules. After that, the classes split. One group studied worked examples, the traditional diet of verb spelling instruction. The other group studied worked examples and, crucially, practised recalling the rules from memory: three spaced retrieval sessions across a week, each requiring pupils to reconstruct the full rule set from memory, then check against a correct version on the board. A week later, both groups sat the same tests.
The first result was what anyone familiar with the retrieval‑practice literature would expect. On a retention test, the group that had practised retrieval recalled the rules at about 88 per cent accuracy. The control group managed about 43 per cent. This is a vast gap, of the sort that researchers have been documenting for two decades.
But the second result was harder to explain away. On an application test, where pupils had to actually spell the verbs correctly in sentences, both groups improved from pre‑test to post‑test, but they performed essentially identically to each other. The retrieval group scored about 53 per cent correct; the worked‑examples group about 57 per cent, a difference that was not statistically reliable. The conundrum here is that the students who knew the rules roughly twice as well did not apply them any better.
So what’s going on?
Two Processes, Not One
This is the puzzle the study leaves on the table. In the ordinary logic of classroom teaching, more memory should mean more capable use; the group with the rules firmly encoded should have at least edged the group without them. That they did not is a small but significant crack in a picture we usually take for granted.
The framework the Dutch researchers invoke to explain the dissociation, Koedinger’s Knowledge‑Learning‑Instruction theory, is a useful modern refinement of Ryle. Acquiring a rule‑based skill, on this account, depends on two separable cognitive processes. One is memory: the rule must be encoded, retained, and retrievable on demand. The other is induction: the learner must abstract the patterns that govern when and how the rule applies, and must integrate the rule into a live problem‑solving procedure.
These processes can be strengthened independently. But they can also fail independently. A learner can have a rule robustly encoded and still fail to summon it at the right moment; a learner can intuit when a rule applies while struggling to recall its exact form. From the outside, both failures look the same. From the inside, they are distinct malfunctions requiring distinct repairs.
The temptation, especially when we are under pressure to demonstrate impact, is to treat both as memory problems and to reach for the same tool. The error is not to underestimate memory. Memory matters enormously. The error is to treat memory as enough. It is the substrate on which skill rests, not the skill itself; the fact that a rule has been written into memory says little, on its own, about whether it will appear on the page when it is needed.
What Retrieval Practice Is For
This places retrieval practice, (probably the most celebrated and yet lethally mutated learning strategy of the last twenty years), in an awkward but instructive light. Retrieval practice does exactly what the research has always said it does: it produces durable memory traces but that is not enough. The Ophuis‑Cox paper replicates this finding with characteristic elegance; the retrieval group’s retention score did not merely edge the control, it doubled it.
The more interesting question I think is when and for what retrieval practice does this job best. A second recent study, this time with undergraduates learning word pairs, explores this further. Gupta, Pan and Rickard asked whether retrieval practice helps equally when the cue and target are strongly related in meaning or only weakly connected. Participants studied word pairs that were either low or high in semantic relatedness, then either restudied them or practised cued recall with feedback, before a delayed test 24 hours later.
As you would expect, testing with feedback beat restudy across the board: retrieval practice still does its basic job. But when they modelled the data using a dual‑memory framework and looked not just at means but at the full distribution of performance, they found that the benefit of testing was about 26 per cent smaller for highly related pairs. The standard ANOVA showed no interaction, yet the model and the cumulative distributions made it clear that the relative efficacy of testing was systematically reduced in the high‑relatedness condition.
A recent review by Roelle, Rummer, Schweppe and colleagues in Learning and Instruction makes the same point at a higher altitude. Under the banner of “meaningful learning”, they distinguish between the work of constructing rich, integrated knowledge and the work of maintaining those outcomes over time. Retrieval practice, in their account, belongs firmly in the second camp: it is a way of stabilising and refreshing knowledge that has already been meaningfully learned, not a shortcut to meaning by itself. On that view, the Dutch verb‑spelling study is not an anomaly; it is what you should expect when a consolidation tool is asked to carry the full burden of construction as well.
What Kind of Knowledge?
So the Dutch pupils could retrieve the verb rules at will on a structured test, but this did not yet buy them better performance in real spelling. The undergraduates benefited from testing on both tightly and loosely related word pairs, but the extra gain from testing was largest where the semantic fabric was thinnest. Taken together, those results suggest that what retrieval practice buys you depends as much on the underlying knowledge structure and how it’s sequenced as the act of retrieval itself.
This matters for how we design curricula and practice. If you are working with isolated facts and arbitrary labels, retrieval practice is probably doing more work than anything else you have. If you are working inside a rich, well‑sequenced knowledge domain, a good part of what looks like a “testing effect” is really the underlying semantic structure doing its job, and the incremental advantage from retrieval will be smaller. And in neither case should you expect practice tests, by themselves, to conjure fluent, procedural performance out of a rule that has barely had time to bed in.
The clearest example here I think is vocabulary. In a curriculum where new words are taught as standalone lists, Friday spelling tests and vocabulary flashcards stripped of context, retrieval practice is the main thing keeping those words in a pupil's mind. Each word is an isolated brick, held up by nothing but rehearsal. In a curriculum where the same words are encountered repeatedly inside a knowledge-rich programme of study, where a child meets "famine" in a text about the potato blight, then in a history lesson on Ireland in the 1840s, then in a poem about harvest, then in a science unit on soil, each encounter does double duty: it rehearses the word and thickens the semantic web the word sits in.
The quiz in the second curriculum looks less heroic because much of the work is being done elsewhere. The quiz in the first curriculum looks essential, because there is nothing else propping the word up. This is the shape of the Hirschian point: knowledge-rich curricula do not make retrieval practice unnecessary; they change what it is doing and how much of the load it is being asked to carry.
The upshot of all this is not that retrieval practice is overrated, but that we must be much more exact about what kind of knowledge we are strengthening, and what we expect that knowledge to support.
The Fluency Threshold
Why might strong memory for the rule fail to convert into fluent application? The Ophuis‑Cox team propose a mechanism worth taking seriously. Applying Dutch verb spelling rules is not a single cognitive act; it is a sequence. The speller must identify the verb, determine whether it is finite, ascertain its tense, analyse the subject, select the appropriate rule from several candidates, and apply it. Each step competes with the others for limited working memory.
If the rule itself is not yet fluent, the act of retrieving it consumes cognitive resources that are needed downstream for the grammatical analysis. The speller runs out of room before she reaches the answer. Under pressure, the brain does what brains do under pressure: it defaults to the fast, frequency‑based system. It picks whichever ending it has seen most often and hopes for the best, exactly the pattern spelling researchers have reported in Dutch and other languages.
This is why fluency is not optional. The point of practising a rule to the point of automaticity is not aesthetic, and it is certainly not anti‑intellectual. It is that a rule which cannot be applied without effort cannot be applied at all, once the other cognitive demands arrive. Fluency frees the working memory that application requires. Without it, knowing the rule and not knowing the rule produce the same output: whatever the frequency‑weighted lexicon happens to throw up.
And, as Gupta and colleagues suggest, fluency is not just a matter of more practice; it depends on how well the new material is anchored in what is already known. Highly related items exploited pre‑existing semantic scaffolding and showed smaller marginal gains from testing than weakly related items. What looks like “more retrieval practice” in the abstract is in reality different work, depending on the underlying knowledge structure.
On this reading, the Dutch study did not fail to help children apply the rules. It succeeded in helping them encode the rules, and ran out of time before it could build the fluency that application requires. Three sessions across a week is a respectable dose of retrieval practice by the standards of the field. But it is nothing like enough to push a freshly acquired procedure into the zone where it costs the working memory nothing to deploy.
Fine-Tuning Retrieval Practice
If this diagnosis is right, four consequences follow for how we think about practice.
The first is that the criterial task matters. If you want pupils to apply a rule fluently, you must eventually practise applying the rule fluently; it is not enough to practise just remembering what the rule is. The Dutch children rehearsed retention and got better at retention. They did not rehearse application at scale, and application did not improve beyond what the worked examples had already produced. What you practise is what you get.
The second is that sequence matters. Retrieval practice is probably best understood as one stage in a longer pipeline, not as a standalone intervention. Get the rule into memory first, yes; then build fluency with it; then practise applying it under the conditions in which it will be used. Each stage has its own instructional signature. Skipping stages produces exactly the dissociation the Ophuis‑Cox study observed.
The third is that dosage matters. Three practice sessions is a generous science‑of‑learning intervention. But it’s a trivial amount of practice by the standards of actual skill acquisition. A child who has rehearsed the times tables three times in a week does not yet know the times tables; she has met them. Much educational research operates at a dosage far below the level at which the interesting thresholds appear, and then reports findings that would probably shift if the intervention ran for another month.
Curricular architecture matters. If retrieval practice yields its largest marginal benefits precisely where semantic relatedness is low, then we should not be surprised if quizzes look most heroic in knowledge‑thin curricula, and more modest in content‑rich ones. Where we have failed to do the groundwork of building a shared store of background knowledge and vocabulary, retrieval practice ends up doing double duty as both memory booster and makeshift semantic scaffolding. Where that groundwork has been done, tests are still useful, but they are nudging an already coherent structure rather than trying to hold a pile of sand together.
For example, consider a child who can chant “i before e except after c” on demand, reliably, word-perfect, and who still writes “recieve” three times a week. She knows the rule. She has practised the knowing. What she has not practised is the act of summoning the rule in the fraction of a second when her hand is choosing between two letters. The gap is not a gap in knowledge. It is a gap in what the knowledge has been rehearsed for.
The test of a rule is not whether a child can produce it, but whether she can deploy it under the cognitive conditions in which it is needed, and producing and deploying are not the same exercise
Curricular architecture: thin vs rich knowledge
And when you zoom out from lessons to whole curricula, the shape of the knowledge you are teaching becomes the next constraint. Retrieval practice does not land on a blank canvas; it lands on whatever semantic structure the curriculum has already built. Two quick contrasts make the point.
Knowledge‑thin example
A Year 7 geography scheme is built around a set of skills (“interpret graphs”, “evaluate sources”) with little substantive content. Retrieval practice here often means quizzing on isolated terms (“delta”, “urbanisation”) or procedural phrases (“PEEL paragraph”). In that setting, quizzing has to act as makeshift semantic scaffolding: it is the only thing holding those stray terms in memory long enough to be used on the end‑of‑topic test. Unsurprisingly, the test scores look “transformational” when quizzes are introduced.Knowledge‑rich example
A Year 7 history curriculum spends weeks building a dense web of knowledge about, say, the English Reformation: who the main actors were, what they believed, how events unfolded, how it links to Europe. Pupils read multiple narrative accounts, work with timelines and maps, and see the same ideas in different guises. In that context, weekly retrieval quizzes – dates, names, cause‑and‑effect links – still help, but the material was already semantically connected. The marginal benefit of quizzing is smaller, because the curriculum itself is doing much of the work that quizzes had to do in the barren geography scheme.
Conclusion
Put crudely, retrieval practice is at its most potent when the connections are weak and need to be built, not when the material is already woven tightly into a semantic web. It is a strategy of the first kind in two senses: it strengthens memory rather than induction, and within memory it is especially good at shoring up the lonely, fragile facts that would otherwise fall away. The error in the field, if there has been one, has been to treat retrieval practice as though it were a general‑purpose solvent for the whole of learning. It isn’t.
These new papers highlight to me that Ryle’s distinction still does not have the traction it deserves. Education continues to assess what is easy to assess, to practise what is easy to practise, and to treat the resulting scores as though they measured the thing that mattered. The Dutch children could recite the rule beautifully. When it came to the thing the rule was actually for, they performed no better than the children who could not recite it at all.
If the point of teaching the rule was the spelling, we were not teaching the thing we thought we were teaching. If the point of quizzing the facts was to enable fluent reasoning, we were again stopping too early, at the point where knowing‑that feels satisfying and measurable. The child who knows the rule and still writes the wrong form is not failing. She is telling us, with the precision of a diagnostic instrument, that we stopped too early.
Mohan W. Gupta, Steven C. Pan & Timothy C. Rickard, ‘Semantic relatedness and the efficacy of retrieval practice’, npj Science of Learning (2026).
Fieke H. A. Ophuis‑Cox, Leen Catrysse & Gino Camp, ‘Mixed effects of retrieval practice on the acquisition of procedural knowledge in verb spelling education’, Learning and Instruction (2026).





Great post.
I recently came back into the world of teaching after several decades in practice/industry. The world has changed a bit since I last taught, and I am trying to find my way back in, to find ways to be pedagogically effective. I have found your posts to be helpful in figuring out what works and what does not.
Your comments in this post made me wonder if I am more focused in the classroom on the recall component than I should be. It made me question whether, in the finance and business statistics undergraduate foundational classes I teach, I am really helping students get good at recalling concepts and formulas, or whether I am in fact getting them to a place where they can apply them. I suspect the practice problems and homework assignments help, but I am not convinced I am creating enough circumstances where students have to work through problems not directly addressed in lectures or discussions.
I am hoping to use the summer to sort through what this means for my classroom in the Fall — how to implement these ideas without overburdening my students or scaring them away!
Thank you for the post and the substack!
Reminds me of the anecdote of the toddler singing the Mickey Mouse clubhouse song. They belt out M-I-C-K-E-Y M-O-U-S-E at the top of their voice but when they are asked to spell Mickey Mouse they haven’t a breeze!!