SUS Score Calculator: Grade Your Usability Scale Result
Enter each participant's ten answers and this returns the System Usability Scale score, the letter grade it earns, and how far the study average sits from the 68 that counts as normal. The number that trips people up is 70: it looks like a pass mark, but SUS is not a percentage, and 70 is an ordinary middling result rather than a good one.
Score a System Usability Scale questionnaire
Answer as one participant did, from 1 (strongly disagree) to 5 (strongly agree), then bank the response and move on to the next person. The average, grade and interval update every time you add or remove someone.
Responses banked
How the scoring works
The System Usability Scale has ten statements that alternate between positive and negative wording, each answered on a five point agreement scale. Every answer converts to a 0 to 4 score: for the odd numbered statements you subtract 1 from the answer, and for the even numbered ones you subtract the answer from 5. Add the ten converted values, which gives 0 to 40, then multiply by 2.5 to land on the familiar 0 to 100 range.
That multiplication is the reason so many teams misread their result. Scaling to 100 makes the output look like a percentage, so a 72 gets reported as a comfortable pass when it is nothing of the sort. Nobody scored 72 per cent of anything: the score is a position on a scale whose meaning comes entirely from comparison with other products.
The alternating wording is deliberate. Answering a run of ten agreeable statements invites straight-lining, where a participant ticks the same column all the way down without reading. Flipping every other statement means a straight-lined form produces a score of exactly 50, which is a useful smell test when you are entering responses.
What your number actually means
The reference point most practitioners use is 68, the average across a large database of SUS studies compiled by Jeff Sauro. Scores above it are above average, scores below it are below average, and the distribution is tight enough that small differences matter more than they look. The letter grades in the calculator come from the curved grading scale that Sauro and James Lewis derived from that same body of studies, which is why a 70 lands as a C rather than the near miss it appears to be, and why the A bands start around 80.
Two habits make the number more honest. First, treat SUS as a relative measure: your own previous release, a competitor's product tested with the same script, or a benchmark of similar tools are all better comparisons than an abstract target. Second, report the interval alongside the mean. A study of eight people can easily produce an interval more than twenty points wide, and a redesign that moves the mean from 71 to 76 inside intervals that wide has not been shown to have moved anything.
How many participants you need
SUS is a questionnaire, not a usability test, and the sample sizes differ. Five or six participants will surface most of the serious usability problems in a session, but they will not pin down a score: individual SUS responses scatter widely, so a handful of people gives you a very wide interval. If the score itself is going to be used for a decision, plan for something closer to twenty responses, and if you are comparing two designs, size each group so the interval is narrower than the difference you care about.
The interval in the result uses the t distribution on your entered scores, which is the standard approach for a mean from a small sample. It answers one question: given this spread, where is the real average likely to sit? It cannot tell you whether the sample is representative. Ten responses from your own staff will produce a tidy interval around a meaningless number.
Running it so the score means something
- Ask straight after the session, before the debrief. Once you have discussed the problems, participants recalibrate and their ratings drift.
- Do not reword the statements between rounds. Swapping "system" for your product name is fine and common, but do it consistently, because comparing scores from two differently worded questionnaires is comparing two different instruments.
- Push for a complete form. A missing answer breaks the arithmetic; the convention when someone genuinely cannot answer is to mark the centre point, which contributes 2 of the 4 possible points.
- Keep the tasks constant. The score reflects what people were asked to do. An easier task set lifts the score without a single change to the interface.
- Read the per statement rows, not just the total. Two products can both score 70 while one is confusing and the other is merely slow to learn, and the weakest statement usually points at which.
What SUS will not tell you
It measures perceived usability and nothing else. It says nothing about whether people completed the tasks, how long they took, whether they would pay, or how the product compares on visual design. A product can score well because it is simple and still fail commercially because it is missing the feature people came for. Pair the score with task success and time on task from the same sessions, and the three together will usually explain each other.
If you are setting up the sessions that feed this, the usability testing guide covers recruitment, task writing and moderation, and what user research is sets out where questionnaires sit next to the other methods. For a heuristic pass before you put anyone in front of the product, use the UX audit guide. When you have a fix worth proving with live traffic rather than a lab, size it with the sample size calculator, and read the confidence interval calculator for the same interval logic applied to conversion rates.
More from Experimento
related resultsBest A/B Testing Tools in 2026, Compared by What They Actually Do
A practical 2026 comparison of the best A/B testing tools by what each one is actually good at, from Optimizely and VWO to GrowthBook and PostHog.
read result →A/B Testing Tools Compared: Which Platform Fits Your Team
A/B testing tools compared for 2026: VWO, Optimizely, AB Tasty, Convert, GrowthBook, Statsig and PostHog, matched to marketing, CRO and engineering teams.
read result →Data-Driven Design: Let Research Guide Your Product Decisions
A practical guide to data-driven design: how to pair analytics with user research, run honest tests, and make product decisions on evidence, not opinion.
read result →