The Monthly Dispatch - What’s New in Learning Science - May 2026
New evidence on the limits of retrieval practice, phone bans that don't move test scores, ChatGPT as cognitive crutch, and why early vocabulary growth tells us surprisingly little about later reading.
Memory, Learning, and Knowledge
1. Does interleaving still work when the content gets hard?
This study tested a fairly important boundary condition in the “desirable difficulties” literature. Interleaving is often promoted as a powerful learning strategy, but most of that evidence comes from relatively simple materials. Here, 376 upper secondary students learned complex physics content under four conditions: interleaved vs blocked practice, and collaborative vs individual learning.
The key result: interleaving alone did not improve learning and actually led to worse long term outcomes. The benefit only appeared when interleaving was combined with collaboration, and that advantage held even after eight weeks. This study fits neatly into a growing body of work showing that desirable difficulties are conditional, not universal. It aligns with research showing interleaving benefits are strongest in lower element interactivity tasks (for example category learning or maths problem types) and weaker or negative in high complexity domains. It also connects with key ideas from John Sweller, particularly the aforementioned element interactivity and the idea that instructional techniques must match task complexity.
2. Is spacing effective because of memory or because of the teacher?
This study takes a familiar finding the spacing effect and tries to open the black box of why it works in classroom like settings. Students learned psychology concepts via instructional videos in either a spaced or massed format. The spaced group performed better on learning outcomes, which is entirely consistent with decades of research. The twist is that the authors also measured interpersonal brain synchronisation between teacher and student using EEG, and found that spacing increased this synchrony.
The key claim is not just that spacing improves learning, but that it may do so partly through social and perceptual mechanisms. Students in the spaced condition reported more positive views of teaching quality and interaction, and these perceptions mediated the relationship between brain synchrony and learning outcomes.
3. Retrieval practice works well for remembering things. Whether it helps students actually use what they remember is a different question entirely.
Fieke Ophuis-Cox’s doctoral dissertation at the Open University of the Netherlands reports four classroom studies conducted with real primary school pupils using existing teaching materials, not lab-based tasks with undergraduates. The headline finding across the programme is clear: retrieval practice reliably improves factual retention across domains, age groups, and settings. But the story gets considerably more interesting when the outcome is comprehension or application rather than recall.
In Study 1, second-graders practising multiplication facts with spaced flashcards outperformed a chanting-and-restudy group on both accuracy and fluency, at both five-minute and one-week delays. So far, so textbook. In Study 2, fifth- and sixth-graders studying expository texts showed a different pattern: retrieval practice improved retention of text content more than summarisation or restudy, but summarisation produced better comprehension outcomes. Retrieval practice did not outperform restudy on comprehension. In Study 3, fourth-graders learning verb spelling rules saw the same split: combining retrieval practice with worked examples substantially improved long-term retention of the rules, but this did not translate into better rule application. Students could recite the rules but still misspelled the verbs.
I’ve written more about this here:
Instruction and Practice
4. Do talking heads in video lessons actually help learning or quietly undermine it?
This study finds that showing the instructor’s face in video lectures often distracts attention and can harm learning, with its effects depending heavily on the teacher’s particular style.
Using an eye tracking experiment with 181 university students, the researchers compared videos with and without a visible instructor across three teaching styles: autonomy supportive, controlling, and neutral. The results are not flattering for the “talking head” trend. Instructor presence consistently pulled attention away from the learning material itself. Worse, under autonomy supportive teaching it reduced retention and concentration, and under controlling teaching it increased cognitive load and reduced concentration. Only under a neutral teaching style did instructor presence show a modest benefit for retention.
5. Directed inquiry eliminated student misconceptions in biology. The lecture-based approach made them worse.
This quasi-experimental study compared directed inquiry activities against a transmissive (lecture-based) approach for teaching primary school students about human organ systems. The directed inquiry group achieved a significantly higher frequency of correct scientific concepts after instruction, and — more strikingly — their misconceptions were eliminated. In the transmissive group, the frequency of misconceptions actually increased after instruction.
A few caveats before drawing strong conclusions. “Directed inquiry” here is not the same as unguided discovery. These were structured, teacher-led activities with clear learning goals, so closer to guided exploration than to the kind of minimally guided instruction that we know is not effective. And the finding that didactic teaching entrenched misconceptions may tell us more about the quality of that particular teaching than about lectures in general. If the transmissive instruction did not surface and directly address pupils’ existing mental models, it would have left those models intact or even reinforced them through superficial coverage. The study is a useful reminder that instruction which doesn’t engage with what students already believe risks talking past them entirely, but it should not be read as a blanket indictment of explicit teaching, which is by its very nature, highly interactive.
6. Worked examples help elementary students with maths, but adding self-explanation prompts may not always add value, and could even complicate things.
This study from the MathByExample research programme examined how worked examples paired with self-explanation prompts affect elementary students’ mathematics motivation and perceptions of errors. It sits within a larger body of work from this team: their 2023 meta-analysis found a medium effect (g = 0.48) for worked examples on maths performance, but also found that self-explanation prompts as an added feature yielded a negative moderating effect relative to worked examples without prompts. That counterintuitive finding — that asking students to explain their reasoning during worked examples can actually reduce their effectiveness — likely reflects the additional cognitive load that self-explanation imposes on novice learners still building basic schemas.
The current study shifts focus from pure achievement to motivation and error perception, asking whether error-focused worked examples change how students perceive mistakes. For teachers, the broader programme message is important: worked examples remain one of the most reliable instructional tools we have for building initial competence, but layering additional demands on top of them needs to be done carefully and with attention to where the learner actually is. More is not always more.
Reading and Language Development
7. Surprisingly, children with strong vocabularies at age four made the least progress in reading comprehension over the following decade.
This one will surprise many reading advocates. Using sequential latent growth models with two large representative Australian samples (N = 4,983 and N = 4,570), the researchers tracked receptive vocabulary development from age four to eight and reading comprehension from age eight to fourteen. The cross-sectional picture looked exactly as you would expect: stronger vocabulary at age four positively predicted reading comprehension at age eight. But the growth picture told a completely different story. Children who started with stronger receptive vocabularies at age four showed less growth in reading comprehension over the following years. And growth in vocabulary itself was unrelated to growth in reading comprehension from age eight to fourteen.
Before anyone takes this as evidence that vocabulary instruction doesn’t matter, the finding is almost certainly a ceiling and regression-to-mean effect: children who start high have less room to grow, and extreme scores naturally regress. But it is a sharp corrective to the assumption that early vocabulary advantage automatically compounds into accelerating reading gains. It doesn’t. It suggests that whatever drives ongoing reading comprehension growth beyond the early years — likely domain knowledge, text structure awareness, inferencing, and metacognition — is partly independent of the lexical foundation that gets children started. Vocabulary is necessary but not sufficient, and measuring it at age four tells you where a child starts, not where they’re headed. Ultimately, vocabulary is a strong predictor of where children start in reading comprehension, but a weak predictor of how quickly they improve over time.
8. AI-scaffolded reading strategy instruction outperformed both teacher-led instruction and business-as-usual, but the design makes it hard to say exactly why.
In this quasi-experimental study, sixty Jordanian undergraduates learning English as a foreign language were divided across three conditions: AI-assisted instruction (teacher-mediated with ChatGPT scaffolding), teacher-led strategy instruction, and business-as-usual. After sixteen weeks, the AI-assisted group outperformed both others on reading comprehension (partial η² = .43, a large effect) and showed the largest gains in metacognitive awareness of reading strategies (d = 1.49). The teacher-led group also outperformed business-as-usual, reinforcing that explicit strategy instruction matters regardless of medium.
The caveat, which the authors themselves flag clearly, is that each condition was a single intact classroom section with no random assignment of students. More importantly, the AI-assisted condition bundled multiple elements — ChatGPT scaffolding, weekly AI tasks, and continued teacher mediation — making it impossible to determine which specific feature drove the advantage. Was it the AI? The extra practice? The novelty? The additional task structure? The study tells us that a well-designed instructional package that includes AI can outperform one that doesn’t, but it does not tell us that AI was the active ingredient. For teachers, the transferable insight is that explicit reading strategy instruction reliably beats leaving strategy use implicit, and that AI may be a useful tool within such instruction — not a replacement for it.
Motivation, Self-Regulation, and Wellbeing
9. Students collaborating on scientific texts gradually developed more sophisticated ways of regulating each other’s thinking, but this was a painstaking, slow process.
This microgenetic study tracked two pairs of ninth-graders across twelve weekly sessions of collaborative document-based inquiry in biology. Students read texts making conflicting scientific claims with varying credibility and had to coordinate multiple sources. The researchers coded every conversational exchange for self-regulation, co-regulation, and shared regulation of epistemic thinking — that is, how students managed their goals, standards, and reasoning processes when evaluating knowledge claims.
The key finding was that shared regulation (where both students jointly managed their epistemic approach) was the most common form, and it increased over time, along with regulation of epistemic ideals; students gradually raised their standards for what counted as good evidence. This is encouraging, but the study’s design, two purposefully selected dyads, no control group, no achievement outcomes, means it demonstrates that social epistemic regulation develops, not how much it matters for learning. For teachers interested in collaborative inquiry with multiple sources, the implication is that students can develop increasingly sophisticated evaluative habits, but it takes sustained practice over weeks, not a single lesson on “critical thinking.”
Classrooms, Behaviour, and Attention
10. Changing how students are grouped had almost no impact on attainment.
This is an important one. A large-scale trial from the Education Endowment Foundation examined whether altering student grouping practices which we call “setting” in the UK, including reducing attainment grouping and increasing mixed-attainment teaching, improves outcomes for secondary school pupils. The study involved a substantial number of schools and used a randomised design, making it one of the more robust pieces of evidence in this space.
The headline finding is straightforward but will be very controversial: there was no meaningful impact on overall attainment. Pupils in schools that changed their grouping practices performed no better, on average, than those in control schools. There were also no consistent improvements in attitudes to learning or related outcomes. In short, reorganising how students are grouped did not translate into measurable gains.
This challenges a long-standing assumption in education that grouping structures are a major lever for improving outcomes. Indeed, some education academics have referred to setting or streaming pupils as “symbolic violence”. The mechanism often proposed is that mixed-attainment teaching raises expectations and improves access to high-quality instruction, particularly for lower-attaining pupils. That may be true in principle, but this study suggests that changing grouping alone is not sufficient. Without changes to curriculum, instruction, or teacher practice, the structural shift does very little.
Becky Allen has written about this very well here.
11. The largest study of school phone bans ever conducted finds they dramatically reduce phone use but have essentially zero effect on test scores.
This is one of the most controversial findings of the year so far. Hunt Allcott, Thomas Dee, Angela Duckworth, Matthew Gentzkow and colleagues analysed data from nearly five thousand US schools that adopted Yondr lockable phone pouches, combining GPS tracking data from 35 million phones, state standardised test scores, student wellbeing surveys, and two original national surveys. The research design — staggered difference-in-differences with pre-registration — is about as rigorous as observational education research gets.
Their findings are clear and uncomfortable for both sides of the phone debate. Phone pouches dramatically reduced in-school phone use: GPS pings on campus dropped roughly 30%, and teachers reported student phone use for personal reasons falling from 61% to 13%. But the average effect on test scores was precisely zero. The researchers can rule out improvements larger than 0.008 student-level standard deviations, in other words, not “we couldn’t detect an effect,” but “we can confidently say the effect is negligibly small.” High schools showed a modest positive effect in maths (about 0.024 SD, roughly equivalent to one-fifth of the effect of having a high value-added teacher), while middle schools showed small negative effects.
For educators, this should temper expectations in both directions. Phone bans appear to be a genuine wellbeing intervention — but not a learning intervention, at least not one detectable at the level of standardised test scores. This is an important distinction. The phone-free classroom may be calmer, less distracted, and better for students’ psychological health without necessarily translating into measurable academic gains. Schools implementing phone policies should do so for social and developmental reasons, not because they expect results day scores to move. And the short-term disruption is real — institutions need to plan for it rather than being blindsided when behaviour gets worse before it gets better.
SEN, Individual Differences and Inclusion
12. An AI-powered voice-first learning system designed specifically for visually impaired students achieves excellent usability scores, but hasn’t yet been tested on actual learning.
This design science study presents BlindLearn, a voice-first learning framework built from scratch for visually impaired students rather than retrofitting visual platforms with screen readers. The system uses wav2vec 2.0 for speech recognition and Tacotron 2 for text-to-speech, structured around a five-stage pedagogical model (Audio Activation, Narrative Input, Conversational Elaboration, Voice Practice, Adaptive Feedback). Testing with fifteen visually impaired students aged ten to sixteen in Indonesia produced a System Usability Scale score of 84.3, rated “Excellent” and well above published baselines for assistive technology. Expert validation was also strong, with a mean Content Validity Ratio of 0.89.
The study’s theoretical grounding in Universal Design for Learning and cognitive load theory is appropriate — the modality effect genuinely supports exclusive auditory channel use for this population. The limitation is straightforward: this is a usability study, not a learning outcomes study. The system scored well on ease of use, learnability, and satisfaction, but whether students actually learn more through it than through existing tools remains untested. The authors acknowledge this and recommend a twelve-week RCT with at least sixty participants as future work. The principle — that accessibility should be a design paradigm from the start, not a retrofit — is sound and important.
AI, Technology, and Cognitive Tools
13. Students who studied with ChatGPT felt like they had learned more. Forty-five days later, they had learned less.
This randomised controlled trial with 120 Brazilian undergraduates compared unrestricted ChatGPT use during study preparation against traditional methods (textbooks, articles, databases). Both groups prepared ten-minute peer presentations on AI and machine learning topics, then took a surprise retention test forty-five days later. The ChatGPT group scored 57.5% against 68.5% for the traditional group — an eleven percentage-point gap with a medium-to-large effect size of d = 0.68. The ChatGPT group also spent 45% less time studying (3.2 versus 5.8 hours), and crucially, the retention deficit persisted even after controlling for study time. Prior familiarity with AI tools offered no protection.
The author calls this “borrowed competence” — a fluency illusion where the AI’s articulacy is mistaken for the learner’s own understanding. This aligns directly with what we know about desirable difficulties: learning that feels easy rarely lasts, and removing the productive struggle from encoding undermines consolidation. For educators, the recommendation is not to ban ChatGPT but to sequence its use. Let students attempt their own encoding first, then bring AI in as a retrieval coach or a challenge partner rather than a first-contact information source. The moment AI replaces the thinking, it replaces the learning.
Another related study found that, 14 to 15 year olds using ChatGPT for science inquiry often failed to ask good questions, failed to spot weak answers, and stopped too soon, leaving performance only moderate despite open access to AI.
14. Not all AI feedback is created equal. Metacognitive feedback produces transfer; affective feedback does not.
This brain-imaging study with 87 Chinese college students compared three types of AI chatbot feedback while learning about the cardiovascular system: metacognitive feedback (prompting students to reflect on their understanding), affective feedback (encouraging and motivational), and neutral feedback. Using functional near-infrared spectroscopy alongside learning measures, the researchers found that metacognitive feedback produced significantly better retention and, crucially, transfer to new questions — the latter being the harder test. The metacognitive group outperformed the neutral group on transfer (M = 9.38 vs 6.72) and also outperformed the affective group, which showed no transfer advantage over neutral. Different feedback types activated distinct brain regions, with metacognitive feedback increasing frontopolar and middle temporal activation correlated with metacognitive sensitivity.
For anyone designing AI tutoring systems, this has a direct practical implication: building a chatbot that says “great job, keep going!” is not the same as building one that says “explain back to me what you just learned and where you’re still unsure.” The affective feedback felt supportive but didn’t produce durable learning beyond what students would have achieved anyway. Metacognitive prompting — pushing students to monitor and evaluate their own understanding — was the active ingredient. This aligns with decades of research on self-explanation and elaborative interrogation, and suggests that the most effective AI feedback systems will be the ones that make students think harder, not feel better.
15. A randomised controlled trial gave secondary school students GenAI support for self-regulated learning. It made almost no difference.
This well-designed RCT, following CONSORT standards, assigned 371 German secondary students (Grades 7–9) to one of three conditions across six physics and English lessons: GenAI-supported utility value reflection, GenAI-supported cognitive learning strategy prompting, or standard ChatGPT use as a control. The researchers measured self-reported motivation, elaboration strategy use, interest, effort, and domain-specific knowledge before and after.
The results are almost uniformly null. There was no significant effect on elaboration strategy use (self-reported or tested), effort, interest, or — the one that matters most — domain-specific knowledge. The one exception: the utility value condition helped students maintain their sense of the material’s relevance, while the cognitive strategy condition actually decreased perceived utility value. In other words, prompting students to use specific cognitive strategies with AI made the learning feel less worthwhile to them, not more.
This study matters because it tests exactly the kind of intervention that AI advocates routinely propose: use GenAI to scaffold self-regulated learning. The intervention was carefully designed, the RCT was rigorous, and the answer was: it didn’t work. Not because the AI was bad, but because the bottleneck in self-regulated learning is probably not something a chatbot can address in six lessons. The one bright spot — an exploratory finding that students who engaged more meaningfully with the AI maintained more interest — suggests that quality of interaction matters, but the study was not designed to engineer that quality. For educators considering AI-supported SRL tools, this is a reality check: scaffolding self-regulation is hard, slow, and relationship-dependent. Adding a chatbot to the mix does not change that.
16. Do virtual pedagogical agents actually motivate students? A meta-analysis of 58 studies says: some constructs yes, most no.
This three-level meta-analysis synthesised 304 comparisons from 58 studies to test whether on-screen virtual pedagogical agents improve student motivation. The results are tellingly selective. Agents significantly boosted self-efficacy (g = 0.29) and interest (g = 0.30), but had no significant effect on engagement (g = −0.02), intrinsic motivation (g = 0.00), value beliefs (g = 0.09), or utility beliefs. The “general motivation” composite was significant at g = 0.20, but this largely reflected the self-efficacy and interest components.
Design features mattered: agents with facial expressions showed larger effects on both motivation and self-efficacy. But the overall picture is one of modest, selective influence rather than broad motivational enhancement. Pedagogical agents appear to make students feel slightly more capable and slightly more interested, without meaningfully changing whether they find the material valuable, engaging, or intrinsically motivating. For developers and procurement decisions, this deflates the marketing: virtual agents are not motivational game-changers. They are modest tools with specific, limited effects.
17. A gamification study reports a strikingly large effect on academic performance — large enough to warrant scepticism...
This between-subjects experiment with 71 Finnish university students compared a gamified learning management system (using acknowledgment, chance, and progression elements) against the same system without gamification. The gamified group substantially outperformed the control (M = 0.80 vs 0.49), yielding a Cohen’s d of 1.41 — an effect size that would be exceptional for almost any educational intervention.
Before celebrating, the usual caveats apply with extra force here. The sample is small (71 students after outlier removal). The study was a single session. The outcome measure was task performance within the gamified system itself, not a separate knowledge transfer test. An effect size this large in a real educational context would mean gamification is more powerful than reducing class sizes, high-dosage tutoring, and expert teaching combined. It almost certainly isn’t. What it probably shows is that game elements increase immediate task completion and engagement within the platform — a useful but much narrower finding than “gamification improves learning.” Teachers and ed-tech buyers should be wary of studies measuring outcomes inside the intervention rather than outside it.
Resources
IES Teaching Math to Young Children — Professional Learning Toolkit
The US Institute of Education Sciences has released a free, comprehensive professional learning toolkit based on the What Works Clearinghouse practice guide Teaching Math to Young Children. It covers nineteen weeks of structured professional development across four modules — subitising, meaningful counting, comparing magnitudes, and solving basic problems — with approximately twenty hours of module work and eighty-five hours of classroom implementation. Designed for preschool through kindergarten teachers, it includes video, fillable learning journals, classroom activities, and progress monitoring tools. Based on four evidence-based recommendations from a national expert panel. Freely available at ies.ed.gov.
NIFDI Science of Reading Webinar
Dr. Kurt Engelmann recently presented a fascinating talk for the first webinar in the Science of Reading Series, focused on how Direct Instruction builds reading comprehension, not just decoding or fluency. The session covered DI approaches for preschool, kindergarten, ESL/EAL/MLL, and early elementary students, with a focus on the language and comprehension skills that often get neglected in phonics-heavy implementations. Worth watching if you work with early readers or want to see what DI actually looks like when it’s done properly. Watch the recording here.
Housekeeping
I’ll be in Auckland NZ on 22 May, Brisbane on Monday 25th May and in Melbourne on Friday 29th May doing a full day event on Retrieval Practice, Responsive Teaching & Checking for Understanding.
I have written a free guide on responsive teaching for UNESCO available here. Jamie Clark has made a One Pager to accompany it.





From what I can see Carl, the comparison was Yondr pouch schools, to schools with varying cell phone policies, such as backpacks, lockers, etc.
I hope people don’t see this as a comparison between pouches and no restriction policy schools. It’s also a little early in the game to expect significant effects from cell phone restriction.
It does seem to be good pseudo experimental design, I’m going to dig a little deeper into it.
"The mechanism often proposed is that mixed-attainment teaching raises expectations and improves access to high-quality instruction, particularly for lower-attaining pupils. That may be true in principle..." Why would it be true in principle? I'd have thought this was a purely empirical question.