In minutes, not weeks.

Assessment Design

Formative vs Summative Assessment: Designing Measurable Learning Loops

What decades of feedback research say about instrumenting assessment so it changes learning while there is still time to act on it.

Upstack AI ResearchMarch 11, 202610 min read
  • Home
  • Insights
  • Formative vs Summative Assessment: Designing Measurable Learning Loops
Back to all insights

Two verbs hiding inside one noun

Michael Scriven drew the line between formative and summative evaluation in 1967, and the distinction has been misread ever since as a property of the instrument. A quiz is called formative, a final exam is called summative, as though the format decided the category. It does not. The same task can serve either purpose depending on what happens to the result. If a score is recorded and reported, it is summative. If the information travels back to the learner or the teacher in time to change what happens next, it is formative.

Robert Stake's often-quoted gloss captures it: when the cook tastes the soup, that is formative; when the guest tastes the soup, that is summative. The verb, not the noun, does the work. Assessment becomes formative only when it closes a loop, and most of what gets labelled formative in practice never closes one.

Why the loop is the unit of analysis

D. Royce Sadler's 1989 account of formative assessment set the terms most of the field still uses. For assessment to improve learning, Sadler argued, three conditions must be met at once: the learner has to hold a clear notion of the standard being aimed for, be able to compare their current performance against that standard, and have access to actions that close the distance between the two. A grade satisfies none of these. It reports the gap without revealing its shape or how to cross it. This is why the design question is never "what should we test" but "what will the result let someone do differently."

0.4–0.7
Effect size range for classroom formative assessment
Black & Wiliam synthesis of 250+ studies, 1998
0.73
Effect size of feedback on achievement
Hattie, Visible Learning ranking
+0.51
Retrieval practice vs. restudying, weighted mean
Adesope et al. meta-analysis, 2017
> 1/3
Feedback interventions that reduced performance
Kluger & DeNisi, Psychological Bulletin, 1996
Evidence

The evidence base, read honestly

Paul Black and Dylan Wiliam's 1998 review, Inside the Black Box, remains the anchor citation. Drawing on more than 250 studies, they reported that strengthening classroom formative assessment produced learning gains with effect sizes between 0.4 and 0.7 — large relative to most schooling interventions, and, notably, largest for lower-attaining students, which narrows attainment gaps rather than widening them. Those numbers have been cited loosely enough to invite skepticism, and the honest reading is narrower than the headline: the range reflects a heterogeneous body of work, and the gains depend entirely on the quality of the feedback loop, not on the act of testing.

John Hattie's synthesis places feedback near the top of his ranked influences at an effect size of roughly 0.73, comfortably above the 0.4 threshold he treats as the zone where an intervention is worth the effort. But Hattie and Timperley's own 2007 paper, The Power of Feedback, is careful about which feedback earns that number. Feedback aimed at the task and the process behind it helps; feedback aimed at the self — praise, grades, comparison to peers — often does not.

The counter-evidence that keeps the field honest

The most disciplined result in this literature is a discouraging one. Kluger and DeNisi's 1996 meta-analysis of 607 effect sizes found that feedback raised performance on average — but that more than a third of feedback interventions lowered it. Feedback that directs attention to the self rather than the task, or that arrives without a usable next step, can measurably harm learning. Any system that promises to generate feedback "at scale" is, by default, scaling something that fails one time in three unless it is designed against that failure mode.

Feedback is not neutral. In Kluger and DeNisi's meta-analysis, more than a third of interventions made performance worse — usually when feedback pointed at the person instead of the work.

Why frequent, low-stakes assessment does more than it looks like it should

There is a second mechanism, independent of feedback, that makes frequent formative checks valuable: the act of retrieval is itself a learning event. Roediger and Karpicke's 2006 experiments showed that students who were tested on material retained substantially more after a delay than students who spent the same time restudying it — even though restudying looked better on an immediate test. Retrieving an answer strengthens the memory more than re-reading the answer does.

The meta-analytic picture is consistent. Adesope, Trevisan and Sundararajan's 2017 synthesis of 118 studies found practice testing outperformed restudying with a weighted mean effect of +0.51, and outperformed passive filler activity by +0.93. Spacing those retrieval events apart in time adds a further gain: Cepeda and colleagues' 2006 review found distributed practice reliably beats massed practice for durable retention. Taken together, these results explain why a weekly low-stakes quiz can outperform a single high-stakes exam covering the same content — not because the quiz is graded, but because it forces repeated, spaced retrieval and creates repeated opportunities to act on the result.

  • Retrieval over recognition: generating an answer beats selecting one for long-term retention.
  • Spacing over massing: the same number of practice attempts, spread out, retains more.
  • Low stakes over high stakes: lower consequences reduce the test anxiety that suppresses performance and makes frequent checks tolerable.

Formative and summative, compared on what actually differs

DimensionFormativeSummative
Primary purposeImprove learning in progressCertify learning after the fact
TimingDuring instruction, frequentEnd of unit, term, or program
StakesLow or noneHigh; contributes to grades or credentials
Who acts on itLearner and teacher, immediatelyInstitution, admissions, employers
Feedback loopClosed — result changes next actionOpen — result is recorded and reported
Design riskFeedback that misfires or never returnsConstruct narrowing; teaching to the test
Design

Instrumenting the loop

Designing a measurable learning loop means deciding, before any item is written, what decision each result will inform and how quickly. Nicol and Macfarlane-Dick's 2006 principles reframe good formative assessment as the scaffolding of self-regulation: the goal is for learners to internalize the standard so they eventually generate their own feedback. That reframing changes what an instrumented loop should capture.

Four things a loop has to make visible

  • The standard. Learners cannot close a gap they cannot see. Exemplars, rubrics and worked criteria have to be legible before the task, not revealed with the grade.
  • The current position. Item-level rather than test-level data. A single percentage tells a learner they are behind; a per-objective breakdown tells them where.
  • The next action. Every result should resolve to something the learner can do — a specific reattempt, a targeted resource, a reworked example — not a number to feel a way about.
  • The trajectory. Spaced retrieval only pays off if the system remembers what was missed and resurfaces it. The loop is longitudinal, not a series of disconnected checks.

This is where software earns its place. Manual formative assessment is powerful but expensive; a teacher tracking twenty objectives across thirty students is doing bookkeeping the machine should carry. LearnLab's assessment engines instrument the loop directly — tagging items to objectives, returning process-level feedback rather than scores, spacing the resurfacing of missed material, and surfacing per-objective trajectories to both learner and instructor. The pedagogy is not new. What is new is being able to run it at a scale where every student gets the loop, not just the ones a teacher has time to notice.

The design test for any formative system: does each result resolve into a specific next action a learner can take? If it resolves only into a number, the loop is open, and the assessment is summative in disguise.

Where this leaves assessment design

Summative assessment is not the villain in this story; certification is a real institutional need and a well-built final exam serves it. The failure mode is quieter: treating every assessment as summative by default, recording results that no one acts on, and calling the accumulated grades "data." That produces measurement without learning.

The research points the other way. Frequent, low-stakes, retrieval-based checks, tied to visible standards and resolving into concrete next actions, are among the better-supported interventions in education — provided the feedback they generate points at the work rather than the worker, and provided it actually returns in time to matter. The instrument is almost incidental. The loop is the intervention.

Transform Your Classroom

Ready to Make Learning
Measurable?

See how LearnLab turns coursework into a measurable loop — automatic grading, conversational tutors, and per-student analytics built for educators who want to teach, not grade.

Our Compliance & Security Standards

GDPR
CCPA
EU AI Act Aligned
Regional Hosting
AES-256
TLS 1.3

Hosted on AWS infrastructure with SOC 2 Type II & ISO 27001 certified data centers — with data residency available across EU, Middle East, and other regions

Across WorkLab and LearnLab, our AI assists — it does not replace — human hiring, grading, and academic decisions. The "EU AI Act Aligned" badge reflects alignment with the Act's principles by design — transparency, human oversight, and documentation — not a certification. Learn more.

Upstack.AIUpstack.AI

Filter the noise. Interview real candidates. One link works anywhere—no ATS migration needed.

Upstack AI FZ-LLC
FOAM2471, Compass Building
Al Shohada Road, AL Hamra Industrial Zone-FZ
Ras Al Khaimah, United Arab Emirates

Microsoft Store
Publisher: UPSTACK AI
Store ID: 9NT2GR4TDZ0G

© 2022 - 2025 Upstack.AI • All Rights Reserved

Powered By

React
TypeScript
Tailwind
Python
AWS
SSL

Last updated: 21/1/2026