Almost every medical device human factors program eventually runs both formative and summative usability studies. On paper, the two look similar: recruit representative users, hand them a device or simulated device, and observe what happens. In practice, they answer fundamentally different questions — and treating them as interchangeable is one of the most common gaps we see when reviewing a usability engineering file.
Formative testing asks "what should we change?"
Formative studies are diagnostic. They happen early and often, usually against prototypes that are still evolving, and their purpose is to surface use difficulties while the design is still cheap to change. A good formative round doesn't just log that a participant made an error — it digs into why: a label that wasn't noticed, an affordance that suggested the wrong action, a mental model mismatch between the device and the user's prior experience with similar products.
- Sample sizes are typically small and iterative — five to eight participants per round is common.
- Think-aloud protocols and open probing are encouraged; you want insight, not just pass/fail data.
- Findings feed directly back into design changes and the use-related risk analysis.
Summative testing asks "does the final design work?"
Summative — or human factors validation — testing is confirmatory. It is run against the final, production-equivalent design, under simulated-use conditions that represent real use environments as closely as possible. Its job is to demonstrate that critical tasks can be performed safely and that residual use-related risk is acceptable.
- Participants must represent the actual user population, including edge cases like low-vision users or first-time users with no training.
- Moderators minimize intervention — the point is to observe unaided performance, not to teach.
- Every use error, close call, and difficulty must be analyzed for root cause and risk significance, not just tallied.
A validation study with a flawless completion rate but no serious analysis of near-misses tells a regulator less than a smaller study with an honest account of what almost went wrong.
Where teams get this wrong
The most common mistake is running a single "usability test" late in development and asking it to serve both purposes. It's usually too polished to generate honest formative insight, and too informal to stand up as validation evidence. The second most common mistake is skipping formative rounds altogether to save schedule — which reliably shows up later as use errors discovered during summative testing, when design changes are far more expensive.
How HF Workspace helps
Because formative and summative studies feed the same traceability thread — user needs, critical tasks, risk controls, and evidence — keeping them in one connected record makes it far easier to show a reviewer exactly which formative finding led to which design mitigation, and which summative result closed it out. That's precisely the structure HF Workspace is built around.