A heuristic evaluation is a review of an interface by a few people who check it against a set of usability rules, called heuristics, and write down every place it breaks one. It is quick and cheap. You do not need users, a lab or a budget, and one afternoon is enough for a single flow such as checkout or sign-up.
Most teams use Nielsen's 10 usability heuristics as the rules. This guide covers the six steps, how to score what you find, a worked example, and where the method stops being useful.
Jakob Nielsen and Rolf Molich formalised the method in 1990. A small group of evaluators each inspect the interface alone, compare it against the heuristics, and list the problems they find. The evaluators are the testers, so there is nobody to recruit or schedule.
It finds likely problems. It does not prove that users struggle. For that you need to watch real people, which is what a usability test does. If you are unsure how many people that takes, see Nielsen's 5-User Rule.
Nielsen's scale runs from 0 to 4. Rate each problem on how often it happens, how badly it hurts, and whether people can get past it.
| Rating | What it means |
|---|---|
| 0 | Not a usability problem |
| 1 | Cosmetic only. Fix it if there is time |
| 2 | Minor. Low priority |
| 3 | Major. Important to fix |
| 4 | Catastrophe. Fix it before release |
Imagine three people reviewing the checkout of a delivery app. This is a made-up example, but the findings are the kind you will see in real products.
| What the evaluator saw | Heuristic it breaks | Severity |
|---|---|---|
| Tapping Pay shows no sign of progress for several seconds | Visibility of system status | 4 |
| A declined card shows the message "Error 402" | Help users recognise, diagnose and recover from errors | 3 |
| The phone field rejects (415) 555 0132 and accepts only 415-555-0132 | Error prevention | 3 |
| Removing an item reloads the cart, with no undo | User control and freedom | 2 |
| The same charge is called "Tip" on one screen and "Driver bonus" on the next | Consistency and standards | 2 |
| A promo banner pushes the Pay button below the fold on small phones | Aesthetic and minimalist design | 2 |
The first finding gets fixed today. The last can wait for the next design pass. Three evaluators will not agree on every number, and that is expected. Settling those differences is the point of the merge step.
Each finding points to a fix, and often to a UX law that explains why it matters. A missing status message is a Doherty Threshold problem: people expect a response within about 400 milliseconds. An unhelpful error is a Design Every State problem. A rejected phone number is a Postel's Law problem, because the form should accept what people reasonably type.
A crowded screen is an Aesthetic-Usability Effect and Density Matches Context problem. Naming the law gives you a reason to put in front of the team, not just an opinion.
Use it early, cheaply and often, then use user tests to confirm the problems that matter most.
Reading about heuristics is the easy part. The skill is looking at a screen and naming what is wrong. In UX Quest, each UX law comes with a real screen to judge. Design Every State is a good place to start for status and error problems, and Postel's Law covers forgiving forms. Each law page says whether its challenge is free or part of Pro.
A heuristic evaluation is a usability review in which a few evaluators inspect an interface against a set of rules, called heuristics, and list every problem they find. Jakob Nielsen and Rolf Molich formalised the method in 1990.
Three to five. One evaluator finds only about a third of the problems, and each extra evaluator adds less than the one before, so more than five is rarely worth it.
It depends on the size of the flow. A single flow such as checkout fits in an afternoon, including the meeting where the evaluators merge their findings.
In a heuristic evaluation, evaluators inspect the interface against rules. Usability testing watches real users try to complete tasks. The first finds likely problems cheaply, and the second shows which ones actually hurt.
A score from 0 to 4 that says how serious a problem is, where 0 is not a problem and 4 is a catastrophe that must be fixed before release. Use it to decide what to fix first.