Usability Testing and Research Synthesis
Usability testing is a research method where you watch real people attempt tasks on a product to find where they struggle. Research synthesis turns those observations, along with other UX research data, into organized findings a team can act on.
itWeb development | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - Usability Testing and Research Synthesis
Usability testing is the moment a real person meets an interface without the map in the designer's head. They try to complete a realistic task while a facilitator watches what happens. This is useful because a room full of colleagues can agree that a flow is clear right up until a participant cannot find the thing everyone else has been pointing at for weeks.
A task scenario supplies the goal without supplying the route. "Buy shoes for less than forty dollars" is a test. "Click Products, then Shoes" is a treasure map with the treasure already circled. Leave out labels and navigation hints, because the participant's own path is the evidence. When they hesitate, backtrack, or choose the wrong route, the interface has made a claim about itself and the participant has politely disagreed.
Qualitative testing asks why a task failed. It usually works with five to eight participants and supports formative changes while the design can still move. Quantitative testing asks how many, how much, or how long. It needs more participants and controlled conditions because it measures performance rather than explaining the source of the trouble. Five people can reveal common problems. They cannot make a statistical sample appear by force of enthusiasm.
A moderated session keeps a facilitator nearby for follow-up questions and neutral think-aloud prompts. An unmoderated session trades that depth for speed and scale. Neither is the superior creature in all habitats. Early, ambiguous prototypes often need someone present; a live product with short, clear tasks can be tested without a human sitting beside every participant.
Then comes research synthesis, the part where recordings, notes, and observations stop being a small mountain range and become decisions. Affinity diagramming groups individual observations into labeled themes. Thematic analysis codes data segments before grouping them. Either way, a useful finding names the design element, preserves evidence, and carries a severity rating based on frequency, impact, and persistence. "Registration was hard" is a weather report. "The register button had low contrast" gives a team somewhere to begin.
Read the Intro for the full map and the Slides for the method choices. Use the Cheatsheet while planning a study, then take the Practice tab into a real small test round. The Quiz checks the distinctions that tend to blur when every note on the board appears to be shouting at once.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://www.nngroup.com/articles/usability-testing-101/
Supports
- Definition of usability testing as a facilitator observing a participant perform tasks on an interface
- The three primary goals: identifying design problems, uncovering improvement opportunities, learning user behavior and preferences
- Task-based nature of sessions and the risk of biased task wording
- Qualitative testing as the more common approach, and its focus on insights/findings versus quantitative metrics
- Five-participant guidance for a qualitative study of one user group
- https://www.nngroup.com/articles/task-scenarios-usability-testing/
Supports
- Definition of a task scenario as a contextualized action request, distinct from a bare instruction
- The "buy shoes for less than $40" versus "purchase orange Nike running shoes" example
- Guidance to make tasks realistic, actionable (not self-reported), and free of interface language/clues
- https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/
Supports
- The N × (1 − L)^n problem-discovery formula, with L approximately 31%
- Five participants uncovering roughly 85% of common usability problems in one user group
- Exceptions requiring larger samples — approximately 20 for quantitative studies, approximately 15 for card sorting
- Guidance to test 3-4 participants per group for two groups, 3 per group for three or more groups
- Recommendation to run several small iterative studies rather than one large study
- https://www.nngroup.com/articles/quant-vs-qual/
Supports
- Qualitative testing sample sizes (5-8) versus quantitative sample sizes (30+)
- Qualitative testing answering "why" questions; quantitative testing answering "how many/how much"
- Qualitative testing under flexible conditions with think-aloud as standard
- Quantitative testing under strictly controlled conditions, often excluding think-aloud
- Qualitative testing serving formative goals; quantitative testing serving summative goals
- https://www.nngroup.com/articles/formative-vs-summative-evaluations/
Supports
- Formative evaluation definition and its role determining what works and why during design
- Summative evaluation definition and its role measuring a finished design against a benchmark
- Formative evaluation occurring repeatedly during design; summative evaluation occurring near or after launch
- https://www.nngroup.com/articles/thinking-aloud-the-1-usability-tool/
Supports
- Definition of the think-aloud protocol as continuous verbalization of thoughts while using an interface
- Think-aloud's value as inexpensive, robust, and usable at any project stage
- Challenges including unnatural continuous narration and participants filtering thoughts to appear competent
- Facilitator practices including demo videos, neutral prompts, and discarding facilitator-biased segments
- https://www.nngroup.com/articles/unmoderated-usability-testing/
Supports
- Definition of unmoderated testing as tool-driven, facilitator-free sessions
- Best-fit use cases: live products, finished prototypes, tasks not dependent on imagination/emotion
- Advantages of speed and scale; disadvantages including no real-time clarification and lower engagement
- Shorter typical session length (~20 minutes) compared to moderated sessions
- https://www.nngroup.com/articles/moderated-remote-usability-test-why/
Supports
- Definition of remote moderated testing with a live facilitator over video conferencing
- Comparable quality to in-person testing at lower cost and with flexible scheduling
- Approximately one-hour session length and real-time follow-up questions
- Recruitment and meeting-software setup as the two main resource-intensive elements
- https://www.nngroup.com/articles/measuring-perceived-usability/
Supports
- Comparison of SUS, NASA-TLX, and SEQ as perceived-usability metrics
- SUS administered once post-test for overall system usability; SEQ administered post-task
- NASA-TLX's six-dimension workload measure and its fit for high-stakes domains rather than typical consumer UX
- https://measuringu.com/sus/
Supports
- SUS as a 10-item, 5-point-response questionnaire created by John Brooke in 1986
- SUS scoring procedure — odd items minus 1, even items 5 minus response, sum, multiply by 2.5
- Industry average SUS score of 68, and percentile-based interpretation (80.3+ as top 10%, 68 near 50th percentile)
- Clarification that a SUS score of 70 is not equivalent to "70%" usability
- https://measuringu.com/seq10/
Supports
- SEQ as a 7-point post-task difficulty rating scale
- Immediate post-task administration to capture fresh task-difficulty perception
- Benchmark average SEQ score of 5.3-5.6 across roughly 400 tasks and 10,000 responses
- https://www.nngroup.com/articles/affinity-diagram/
Supports
- Definition of affinity diagramming as clustering related observations into themes
- The three-step process — generating individual notes, clustering into themes, prioritizing clusters
- Affinity diagramming's role converting raw research observations into structured, actionable synthesis
- https://www.nngroup.com/articles/affinity-diagramming-pitfalls/
Supports
- The three common affinity-diagramming pitfalls — no focus question, keyword matching, groupthink
- Fixes for each pitfall, including a visible focus question, sentence-length cluster labels, and silent clustering
- https://www.nngroup.com/articles/thematic-analysis/
Supports
- Definition of thematic analysis as coding data segments, then grouping codes into themes
- The code-then-group order of operations, contrasted with affinity diagramming's cluster-then-label order
- Thematic analysis as slower but more thorough than affinity diagramming
- https://www.nngroup.com/articles/actionable-usability-findings/
Supports
- Guidance to write specific findings naming the responsible design element, not vague symptoms
- Use of severity ratings (low/medium/high) to help teams prioritize findings
- Guidance to offer recommendations as suggestions rather than final redesigns, and to include positive findings
- https://www.nngroup.com/articles/how-to-rate-the-severity-of-usability-problems/
Supports
- The 0-4 severity rating scale definitions, from "not a usability problem" to "usability catastrophe"
- The three factors combining into a severity rating — frequency, impact, and persistence
- https://www.nngroup.com/articles/recruiting-screening-research-candidates/
Supports
- Definition and purpose of a screener — identifying representative candidates and excluding poor fits
- Guidance to disguise the target trait among unrelated options rather than asking directly
- Guidance to exclude UX/design professionals and habitual paid participants from general studies
- Guidance to place hard exclusion criteria early in a screener
- https://github.com/awesomelistsio/awesome-ux
Supports
- Discovery source identifying Dovetail, Maze, and Lookback as current UX research and usability-testing tools
- https://github.com/ttt30ga/awesome-product-design
Supports
- Discovery source identifying UserTesting and Optimal Workshop as current usability-testing tools, listed under its Research/Testing section
- https://maze.co/
Supports
- Maze as a rapid usability-testing platform for prototypes and live products supporting moderated and unmoderated studies
- https://www.lookback.com/
Supports
- Lookback as an AI-assisted remote research platform for live moderated sessions with real-time note capture
- https://www.usertesting.com/
Supports
- UserTesting as a human-insight platform with a large vetted participant network for fast, real-world feedback
- https://www.optimalworkshop.com/
Supports
- Optimal Workshop as a research platform specializing in card sorting, tree testing, and first-click testing for information architecture
- https://dovetail.com/
Supports
- Dovetail as a research repository and analysis platform for organizing and tagging qualitative research data
- https://www.userinterviews.com/
Supports
- User Interviews as a participant-recruitment platform sourcing research candidates from an external panel
- https://mitpress.mit.edu/9780262050296/protocol-analysis/
Supports
- 1984 publication of Protocol Analysis: Verbal Reports as Data by Ericsson and Simon, a foundation for think-aloud methods
- https://dl.acm.org/doi/10.1145/169059.169166
Supports
- 1993 publication of Nielsen and Landauer's mathematical model for finding usability problems
- https://www.nist.gov/itl/tted/iusr-project-industry-usability-report
Supports
- NIST Industry Usability Reporting project began in 1997
- ANSI/NCITS-354-2001 Common Industry Format approval in 2001
- https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=907849
Supports
- ISO/IEC 25062 Common Industry Format for Usability Test Reports published in 2005
- https://www.iso.org/files/live/sites/isoorg/files/archive/pdf/en/isofocusplus_2012-07.pdf
Supports
- ISO 9241-11:1998 definition of usability
- ISO 9241-210:2010 revision of ISO 13407:1999 for human-centred design processes
- https://doi.org/10.1145/3299096
Supports
- Ethnographic study showing that practitioners collaboratively produce usability findings during and after testing
- Potential candidate problems can be dissipated or ignored in collaborative analysis
- https://dovetail.com/pricing/
Supports
- Dovetail offers a free plan alongside paid plans for research-data organization and analysis
