The familiar pattern: a learner asks a question on Monday, gets a satisfying answer, and on Friday cannot retrieve the idea. We tend to read this as a failure of memory. Six weeks of session logs suggest something different. The route back matters as much as the idea itself.
Key takeaways
- Re-framed recall beats repetition. Changing the surface of a recall prompt raised one-week retention from 41% to 58% across 184 learners.
- The gain concentrates where it matters. Ideas learners described as “almost clear” — recognised but not articulable — improved most, reaching 63%.
- Production, not recognition. A new door into the same room forces the learner to regenerate the idea instead of re-reading it.
- Shipped by default. Recall prompts now re-frame by default, with a small seed so the prompt is never a blank page.
The question we kept seeing
In our early cohorts, retention at one week was lower than we expected. Learners who had asked rich, multi-turn questions would return days later and ask a question that sounded almost identical to one from the previous session, but with a different surface framing. They were not forgetting the idea. They were forgetting that they had already asked it.
This is a familiar finding in the spaced-repetition literature: recognition is not the same as recall, and an answer that was clear in context can become opaque out of context. What surprised us was how much the shape of the next prompt mattered.
The observation surfaced in a place we were not looking for it. We had been reviewing retention data for a separate study on spaced repetition and noticed that learners who came back to a topic after a gap, and who themselves rephrased the original question rather than recognising the tutor's rephrased version, did better on the one-week recall prompt. The effect was small—about 4 percentage points in the first month we looked at it—and we almost dismissed it as noise. What kept us looking was that the effect was monotone across cohorts. Every subgroup we checked showed the same direction, and the magnitude grew as the rephrasings grew more substantively different from the original.
The naive explanation is that rephrasing is itself a form of retrieval practice, and the small lift is just the lift from retrieval. We tested that explanation directly and it is wrong. The lift from retrieval practice, in our data, is about 6 percentage points for a single rephrasing. The lift we were seeing was 17 percentage points for the "almost clear" subgroup—more than the lift retrieval alone could explain. The additional lift was coming from somewhere else.
The somewhere else, we eventually concluded, is that a rephrasing is not a single act. It is a small chain: the learner notices that they have forgotten the original framing, retrieves the idea well enough to put new words around it, and then tests those new words against the original idea for fit. Each step in the chain is a small act of construction, and each construction is, in the language of the literature, an elaboration. Elaborations are the most reliable route to durable retention that cognitive psychology has found. We were not getting the lift because the learner had rephrased. We were getting the lift because the learner had rephrased and then tested the rephrase against the original.
What we changed
We ran two small interventions across 184 active learners over six weeks. In one group, the tutor's recall prompt repeated the original question verbatim. In the other, the tutor re-framed the idea through a different surface—asking the learner to recall the same concept from a new angle, often with a smaller case or a counterexample. Both groups received the same total recall time and the same review interval.
The re-framed prompts moved the needle. One-week retention rose from 41% to 58%. The effect was strongest on the kinds of ideas that learners described as "almost clear"—the ones where recognition was easy but articulation was hard.
The first change is the more visible one. When a learner asks "what is the chain rule?" and then, three days later, asks "what is the chain rule?" again, the old tutor would respond with the same definition in slightly different English. The new tutor responds with a comparison: "here is the chain rule, and here is the product rule—what is the same, and what is different?" The shape of the response is different. The surface is also different, but that is incidental. The shape change is what matters.
We chose comparison as the first new shape for a specific reason. Comparison is one of the most reliable elaborative operations in the literature. When a learner is asked to say what is the same and what is different between two ideas, they are forced to identify the underlying features of each idea, because the comparison will only make sense if the features are explicit. A learner who can compare the chain rule to the product rule has, in effect, built a small model of both. A learner who can only re-state the chain rule has built a small surface for one.
The second change—logging the change in shape rather than the change in wording—was the one that made the measurement possible. Until we made this change, we had been measuring recall by counting word overlap between the original and the rephrased version. That metric is a measure of surface similarity, not of shape similarity. Two definitions that use different words have low surface similarity; a definition and a comparison have low surface similarity too. By the surface metric, we could not tell the difference between a re-shape and a re-word. Once we logged the change in shape, the effect became visible.
One-week retention by week of intervention
The route back is not the route forward. If we want the idea to come back when the learner is alone with a blank page, the path has to feel like one they can walk.
Why we think this works
Two explanations seem plausible, and we have not been able to choose between them with the data so far.
The first is that re-framing forces a kind of transfer: the learner has to recognise the same idea under a new description. Transfer is one of the best predictors of durable learning, and it is exactly the work that recognition-only recall does not do.
The second is more procedural. When a learner returns to the original framing, they often re-read the previous answer to remind themselves of what was said. When they return to a new framing, they have to produce the idea without the script. Production is harder, but it is also what they will need to do the next time they encounter the concept in the wild.
If you want someone to remember an idea next month, give them a new door into the same room.
What this changes in the product
Two small changes shipped as a result. First, recall prompts now re-frame by default. We let learners choose to repeat the original phrasing if they want it, but the default is the new angle. Second, the recall queue now carries a small seed—a hint about the new surface—so the prompt is not a completely blank page.
Neither change is dramatic. Both are the kind of thing that only shows up across weeks of use. The interesting question for us is whether the same effect holds for ideas that were learned under time pressure, or for learners who already feel confident about the concept. We have a follow-up planned for September.
What the data says
Across the six-week intervention, 184 learners completed at least one recall cycle in each condition. The headline numbers:
One-week delayed recall · 184 learners
May – Jul 2026The gap between the first two rows is the average treatment effect. The gap between the last two is the subgroup we are now designing around: ideas that feel familiar but cannot yet be articulated. For those, recognition is easy, recall is hard, and a new door changes the outcome most.
Where the effect concentrates
Splitting the cohort by session type shows the pattern is not uniform. The effect is largest for the “almost clear” middle, real for time-pressured sessions, and close to flat for learners who already felt confident—which is what a production account of the effect would predict.
| Cohort | Verbatim | Re-framed | Difference |
|---|---|---|---|
| All learners (n = 184) | 41% | 58% | +17 pp |
| “Almost clear” ideas | 47% | 63% | +16 pp |
| Sessions under time pressure | 38% | 52% | +14 pp |
| Self-reported confident learners | 44% | 49% | +5 pp |
Table 1 — One-week delayed recall by subgroup · differences are percentage-point gaps
How we measured it
Learners were assigned to one of two recall-prompt conditions for the full six weeks. Both groups received identical review intervals and identical total review time; only the surface of the prompt differed. Retention was measured with a delayed free-recall prompt at the one-week mark, scored blindly by two raters. We controlled for subject, prior session count, and self-reported confidence.
We logged the shape of the re-framing—definition to comparison, worked example to counterexample, theorem to special case—and the effect held in every shape pair, with the smallest effect for surface-only rewording.
The free-recall prompt was a single sentence asking the learner to write down, in their own words, what the original answer had been. We deliberately did not give the learner any cues—no multiple-choice options, no fill-in-the-blank, no hint of the original wording. Free recall is the hardest form of retention test, and it is the one that most cleanly measures what the learner has actually constructed, as opposed to what they can recognise. The two raters scored each response on a four-point rubric: 0 (no relevant content), 1 (some relevant content but the central idea is missing), 2 (the central idea is present but supporting detail is wrong or missing), 3 (the central idea is present and supporting detail is largely correct). A score of 2 or 3 counted as "recalled." The inter-rater κ was 0.78, which is acceptable for this kind of rubric.
The "shape pair" finding is one we want to call out. We tried four different shape changes: definition to comparison, definition to example, definition to worked-problem, and definition to contrast (definition paired with a near-neighbour that the learner was likely to confuse it with). Every shape change produced a positive effect over the verbatim control, but the magnitude varied. The largest effect was for comparison (17 percentage points), and the smallest was for surface-only rewording (4 percentage points). This is the cleanest evidence in our data that the kind of change matters more than the amount of change.
Limitations
Three caveats bound what we can claim. The cohort is self-selected learners on one tutoring platform, not a random sample of students. The six-week window is long enough to see the effect appear and short enough that we cannot yet say whether it survives a month. And the raters, though blinded, scored transcripts produced by the same product—so we treat the subgroup splits as directional until the September replication.
Open questions
Three things are still open. Whether the effect survives to one month; whether it holds for ideas learned under time pressure; and whether learners who already feel confident about a concept still benefit from a new door, or whether re-framing mostly helps the "almost clear" middle. A four-week follow-up is planned for September, and we will publish it with the negative results included.
The subgroups we have looked at are: subject (mathematics, programming, history, economics), prior session count (1–3, 4–10, 11+), self-reported confidence at the time of the original question (low, medium, high), and time-of-day of the original question. The shape-effect direction held in every subgroup. The magnitude varied: the smallest effect was in the high-confidence subgroup (7 percentage points) and the largest was in the medium-confidence subgroup (15 percentage points). We did not expect the low-confidence subgroup to show only a 9-percentage-point effect. Our prior was that the low-confidence subgroup would benefit most, on the theory that they have the most to learn. The data says otherwise. The most plausible interpretation is that a learner who is at "low confidence" is not yet at the threshold where they can attempt a re-shape; the re-shape requires some minimum of confidence to attempt, and below that threshold the re-shape is not actually attempted, even when the tutor offers it. This is a hypothesis, not a finding. The next study will test it directly.
A fourth open question: does the shape effect depend on the kind of shape? We have data on comparison, example, worked-problem, and contrast. We do not have data on analogy, on counter-example, or on a shape we have not yet named. The literature suggests that analogy is a particularly strong elaborative operation, especially for abstract ideas, but we have not yet measured it.
A fifth open question, which we have begun to think about: does the shape effect compound across multiple recall cycles? A learner who re-shapes once and then re-shapes again, with a different shape, might do better than a learner who re-shapes once and then is re-tested verbatim. Or the second re-shape might interfere with the first. We do not yet know.
We also publish these notes to make our decisions legible. If the tutor behaves a particular way during a session, there is usually a note here that explains why. The decision to publish these notes is itself a product decision. We have learned that the learners who use Socrates longest, and who get the most out of it, are the learners who develop a working theory of how the tutor behaves. A learner who understands that the tutor will offer a comparison when they ask the same question twice is a learner who can use that knowledge to ask better questions.
References
- Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4–58.