Lesson 21 · Spacing & interleaving (the science this whole puzzle system is built on)

Why This Path Won't Let You Do 20 Lessons on One Topic in a Row

Stage 4 has covered how practice improves performance (Lesson 18), what’s actually changing (Lesson 19’s chunking), and what practice needs to contain to actually work (Lesson 20’s deliberate practice). This closing lesson covers when and in what order practice should happen — and it’s the direct reason this puzzle path is delivered as a slow trickle of lessons rather than a single document you could read straight through.

Terms (standalone):

  • Massed practice: concentrating all study/practice of one topic into a single session or a short span of time — “cramming.”
  • Spaced (distributed) practice: spreading the same total amount of study across multiple sessions separated by gaps of time (hours, days, or weeks).
  • The spacing effect: the robust, century-old finding (first documented by Hermann Ebbinghaus in the 1880s, replicated in hundreds of studies since) that spaced practice produces better long-term retention than massed practice, for the same total amount of study time — and the advantage is often small or negligible on an immediate test, but grows substantially on a delayed test. Cramming can look just as good as spacing right after studying, which is exactly what makes it a trap — the difference only shows up once time has passed and the crammed memory has faded far faster than the spaced one.
  • Interleaving: practicing multiple related-but-distinct skills or problem types mixed together within a single session, rather than in isolated blocks (all of skill A, then all of skill B — “blocked practice”). Counterintuitively, interleaving typically produces worse performance during practice itself (it’s harder, more error-prone in the moment) but better long-term retention and, notably, better ability to correctly identify which technique applies to which problem — a skill blocked practice doesn’t train at all, since within a block you already know which technique to use.
  • Why both effects share one underlying cause: both spacing and interleaving work by deliberately introducing forgetting and retrieval effort between exposures — connecting directly to Lesson 14’s retrieval-practice principle and Lesson 20’s “desirable difficulty.” A massed or blocked session lets you succeed via short-term familiarity that hasn’t actually been consolidated into durable memory; spacing and interleaving force each subsequent attempt to be a genuine retrieval from longer-term storage, which is the effortful act that actually strengthens the trace.

The puzzle (MCQ above)

Think about why the gap between Student A and B would be small immediately after studying but grow larger a month later — what’s decaying at different rates for each of them?


Part 2 — Why does cramming feel more effective, if it isn’t?

Massed practice/cramming has a persistent reputation as “working,” despite the research. Propose a reason cramming feels productive in the moment even though it under-performs on delayed retention — what’s the student actually experiencing during a massed session that creates that impression, and why is it misleading about what will happen a month later?


Part 3 — Reveal: connect this to the puzzle path’s own design

This scheduled routine deliberately maintains a buffer of unsolved lessons rather than dumping the entire curriculum on you at once, delivers lessons across many separate sessions rather than one sitting, and (via the roadmap.md policy) interleaves four different tracks (fp, bayes, cog, ml) rather than completing one track fully before starting the next. Using this lesson’s concepts, explain specifically which design choice corresponds to spacing and which corresponds to interleaving, and why a system optimized purely for your immediate comprehension-per-lesson (rather than durable, transferable learning) might have made different choices.


Part 4 — Reveal: a testable prediction

Based on the interleaving research specifically (not just spacing), predict one concrete way your performance might look worse in the short term but be better in the long term as a direct result of this path mixing fp, bayes, cog, and ml lessons together instead of completing tracks one at a time. Be specific about what kind of test or task would reveal the advantage that a single-track, blocked version of this same curriculum would not.

Two students each spend 4 total hours studying the same material for a test in two weeks. Student A does one 4-hour massed session the night before. Student B studies in four 1-hour sessions spread across the two weeks. On the test, and especially on a SURPRISE follow-up test a month later, who scores higher, and how confident are researchers in this result?