Usability Testing and Research Synthesis
Usability testing is a research method where you watch real people attempt tasks on a product to find where they struggle. Research synthesis turns those observations, along with other UX research data, into organized findings a team can act on.
itWeb development | OpenSkills.info
Intro
Usability Testing and Research Synthesis
Usability testing is a research method in which a facilitator asks a participant to perform tasks on a product while the facilitator observes the participant's behavior and listens to their feedback. The method exists to answer a question no amount of internal review can settle: can a real person, without your knowledge of the interface, actually use it? Research synthesis is the companion discipline that turns the raw output of usability tests, and of other UX research methods, into organized findings a team can act on.
What a usability test looks for
A usability test aims at three outcomes: identifying design problems, uncovering improvement opportunities, and learning about target users' behavior and preferences. Sessions center on realistic tasks rather than open-ended exploration, because a task gives the participant a concrete goal and gives the observer a concrete point of comparison between intended and actual behavior.
The facilitator writes each task as a scenario, not a bare instruction. An instruction such as "purchase orange running shoes" tells the participant exactly what to search for and biases the session toward a specific product. A scenario such as "buy shoes for less than $40" states a realistic goal without revealing interface labels or navigation paths, so the session reveals whether the participant can find their own way to a working method. Scenarios avoid interface language for the same reason: naming a button or menu item in the task removes the very ambiguity the test exists to detect.
Qualitative and quantitative testing
Usability testing splits into two traditions that answer different questions.
Qualitative testing uses a small number of participants, typically five to eight, and asks "why" questions: why did this task fail, what did the participant expect instead, which words confused them. Sessions run under flexible, adapted-as-needed conditions, and the facilitator commonly asks participants to think aloud. Qualitative testing serves formative evaluation — steering a design while it can still change.
Continue the course
This section is part of the paid course.
See pricing to subscribe, or log in if you already have access.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://www.nngroup.com/articles/usability-testing-101/
Supports
- Definition of usability testing as a facilitator observing a participant perform tasks on an interface
- The three primary goals: identifying design problems, uncovering improvement opportunities, learning user behavior and preferences
- Task-based nature of sessions and the risk of biased task wording
- Qualitative testing as the more common approach, and its focus on insights/findings versus quantitative metrics
- Five-participant guidance for a qualitative study of one user group
- https://www.nngroup.com/articles/task-scenarios-usability-testing/
Supports
- Definition of a task scenario as a contextualized action request, distinct from a bare instruction
- The "buy shoes for less than $40" versus "purchase orange Nike running shoes" example
- Guidance to make tasks realistic, actionable (not self-reported), and free of interface language/clues
- https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/
Supports
- The N × (1 − L)^n problem-discovery formula, with L approximately 31%
- Five participants uncovering roughly 85% of common usability problems in one user group
- Exceptions requiring larger samples — approximately 20 for quantitative studies, approximately 15 for card sorting
- Guidance to test 3-4 participants per group for two groups, 3 per group for three or more groups
- Recommendation to run several small iterative studies rather than one large study
- https://www.nngroup.com/articles/quant-vs-qual/
Supports
- Qualitative testing sample sizes (5-8) versus quantitative sample sizes (30+)
- Qualitative testing answering "why" questions; quantitative testing answering "how many/how much"
- Qualitative testing under flexible conditions with think-aloud as standard
- Quantitative testing under strictly controlled conditions, often excluding think-aloud
- Qualitative testing serving formative goals; quantitative testing serving summative goals
- https://www.nngroup.com/articles/formative-vs-summative-evaluations/
Supports
- Formative evaluation definition and its role determining what works and why during design
- Summative evaluation definition and its role measuring a finished design against a benchmark
- Formative evaluation occurring repeatedly during design; summative evaluation occurring near or after launch
- https://www.nngroup.com/articles/thinking-aloud-the-1-usability-tool/
Supports
- Definition of the think-aloud protocol as continuous verbalization of thoughts while using an interface
- Think-aloud's value as inexpensive, robust, and usable at any project stage
- Challenges including unnatural continuous narration and participants filtering thoughts to appear competent
- Facilitator practices including demo videos, neutral prompts, and discarding facilitator-biased segments
- https://www.nngroup.com/articles/unmoderated-usability-testing/
Supports
- Definition of unmoderated testing as tool-driven, facilitator-free sessions
- Best-fit use cases: live products, finished prototypes, tasks not dependent on imagination/emotion
- Advantages of speed and scale; disadvantages including no real-time clarification and lower engagement
- Shorter typical session length (~20 minutes) compared to moderated sessions
- https://www.nngroup.com/articles/moderated-remote-usability-test-why/
Supports
- Definition of remote moderated testing with a live facilitator over video conferencing
- Comparable quality to in-person testing at lower cost and with flexible scheduling
- Approximately one-hour session length and real-time follow-up questions
- Recruitment and meeting-software setup as the two main resource-intensive elements
- https://www.nngroup.com/articles/measuring-perceived-usability/
Supports
- Comparison of SUS, NASA-TLX, and SEQ as perceived-usability metrics
- SUS administered once post-test for overall system usability; SEQ administered post-task
- NASA-TLX's six-dimension workload measure and its fit for high-stakes domains rather than typical consumer UX
- https://measuringu.com/sus/
Supports
- SUS as a 10-item, 5-point-response questionnaire created by John Brooke in 1986
- SUS scoring procedure — odd items minus 1, even items 5 minus response, sum, multiply by 2.5
- Industry average SUS score of 68, and percentile-based interpretation (80.3+ as top 10%, 68 near 50th percentile)
- Clarification that a SUS score of 70 is not equivalent to "70%" usability
- https://measuringu.com/seq10/
Supports
- SEQ as a 7-point post-task difficulty rating scale
- Immediate post-task administration to capture fresh task-difficulty perception
- Benchmark average SEQ score of 5.3-5.6 across roughly 400 tasks and 10,000 responses
- https://www.nngroup.com/articles/affinity-diagram/
Supports
- Definition of affinity diagramming as clustering related observations into themes
- The three-step process — generating individual notes, clustering into themes, prioritizing clusters
- Affinity diagramming's role converting raw research observations into structured, actionable synthesis
- https://www.nngroup.com/articles/affinity-diagramming-pitfalls/
Supports
- The three common affinity-diagramming pitfalls — no focus question, keyword matching, groupthink
- Fixes for each pitfall, including a visible focus question, sentence-length cluster labels, and silent clustering
- https://www.nngroup.com/articles/thematic-analysis/
Supports
- Definition of thematic analysis as coding data segments, then grouping codes into themes
- The code-then-group order of operations, contrasted with affinity diagramming's cluster-then-label order
- Thematic analysis as slower but more thorough than affinity diagramming
- https://www.nngroup.com/articles/actionable-usability-findings/
Supports
- Guidance to write specific findings naming the responsible design element, not vague symptoms
- Use of severity ratings (low/medium/high) to help teams prioritize findings
- Guidance to offer recommendations as suggestions rather than final redesigns, and to include positive findings
- https://www.nngroup.com/articles/how-to-rate-the-severity-of-usability-problems/
Supports
- The 0-4 severity rating scale definitions, from "not a usability problem" to "usability catastrophe"
- The three factors combining into a severity rating — frequency, impact, and persistence
- https://www.nngroup.com/articles/recruiting-screening-research-candidates/
Supports
- Definition and purpose of a screener — identifying representative candidates and excluding poor fits
- Guidance to disguise the target trait among unrelated options rather than asking directly
- Guidance to exclude UX/design professionals and habitual paid participants from general studies
- Guidance to place hard exclusion criteria early in a screener
- https://github.com/awesomelistsio/awesome-ux
Supports
- Discovery source identifying Dovetail, Maze, and Lookback as current UX research and usability-testing tools
- https://github.com/ttt30ga/awesome-product-design
Supports
- Discovery source identifying UserTesting and Optimal Workshop as current usability-testing tools, listed under its Research/Testing section
- https://maze.co/
Supports
- Maze as a rapid usability-testing platform for prototypes and live products supporting moderated and unmoderated studies
- https://www.lookback.com/
Supports
- Lookback as an AI-assisted remote research platform for live moderated sessions with real-time note capture
- https://www.usertesting.com/
Supports
- UserTesting as a human-insight platform with a large vetted participant network for fast, real-world feedback
- https://www.optimalworkshop.com/
Supports
- Optimal Workshop as a research platform specializing in card sorting, tree testing, and first-click testing for information architecture
- https://dovetail.com/
Supports
- Dovetail as a research repository and analysis platform for organizing and tagging qualitative research data
- https://www.userinterviews.com/
Supports
- User Interviews as a participant-recruitment platform sourcing research candidates from an external panel
