Cost per Usability Issue Found: A Simple ROI Formula

A simple ROI model for human usability testing: divide total session spend by the number and severity of actionable issues found. Includes a worked example, severity weights, and common mistakes to avoid.

By the TestTorch team 7 min read
Cost per Usability Issue Found: A Simple ROI Formula

You ran five human testing sessions, watched the replays, and now have 27 comments in a spreadsheet. The hard question is not whether the sessions felt useful; it is whether the spend produced enough actionable evidence to justify doing it again.

The cleanest starting metric is cost per usability issue found. It tells you how much you paid, on average, for each real product problem uncovered by human testers.

Calculate cost per usability issue found with one basic formula

The basic formula is simple:

Cost per usability issue found = total testing spend ÷ number of unique actionable usability issues

If you spend €290 on testing and find 10 unique actionable issues, your cost is €29 per issue found.

The phrase unique actionable matters. If three testers all struggle to find the pricing page, that is usually one issue with stronger evidence, not three separate issues. If a tester says “I do not like the color blue,” that is feedback, but it may not be actionable unless it connects to a clear task failure, misunderstanding, or measurable hesitation.

What counts as total testing spend when sessions are not your only cost?

Your numerator should include the costs needed to get from “we should test this” to “we have reviewed the findings.” For a lean startup, that usually means paid sessions plus the founder or product team’s review time.

If you use TestTorch human testing sessions, founders can buy sessions from €29, and each session includes a vetted tester, full session replay, and written findings. That makes the direct session cost easy to calculate.

For a fuller ROI view, include internal time too. A €29 session is not really €29 if you spend 45 minutes reviewing the replay, tagging issues, and writing tickets.

Cost itemInclude it?Example
Paid testing sessionsYes4 sessions × €29 = €116
Founder or PM review timeUsually yes2.5 hours × €50/hour = €125
Engineering fix timeNo, track separatelyFixing issues is remediation cost, not discovery cost
Recruiting or coordination timeYes, if you do it yourself1 hour finding testers and scheduling
Tooling or platform feesYes, if paid separatelyRecording, survey, or panel fees

Keep discovery cost and fix cost separate. The metric here evaluates how efficiently testing finds problems, not how expensive your product is to improve.

Use severity weighting so a checkout blocker is not equal to a typo

The basic metric treats every issue equally. That is fine for a quick read, but it can underrate a session that finds one serious activation blocker and overrate a session that finds ten cosmetic annoyances.

A better version is severity-adjusted:

Cost per severity point = total testing spend ÷ total severity points from unique actionable issues

Use a simple scale your team can apply consistently. Do not make it too clever; the point is fast comparison across test rounds.

SeverityPointsUse this whenExample
Blocker5The tester cannot complete a key taskCannot create an account because the verification email is unclear
Major3The task is completed, but with serious confusion or delayTester eventually finds billing settings after checking three unrelated menus
Minor1The issue causes friction but does not threaten task completionButton label is vague, but the tester guesses correctly
Observation0Interesting but not actionable yetTester says the dashboard “feels busy” without a clear task impact

Severity points make the metric harder to distort. A test round with 3 blockers and 2 minor issues should look more valuable than one with 12 low-priority copy comments.

A worked example using €29 sessions and review time

Suppose you test a browser-based onboarding flow before inviting 200 beta users. You buy 4 sessions at €29 each, then spend 2.5 hours reviewing session replays and written findings.

InputAmount
Session spend4 × €29 = €116
Review time2.5 hours × €50/hour = €125
Total testing spend€241
Raw notes captured18
Unique actionable issues after deduping11

Your basic cost per issue is:

€241 ÷ 11 = €21.91 per actionable usability issue found

Now apply severity. The 11 issues include 2 blockers, 4 major issues, and 5 minor issues.

Issue typeCountPoints eachTotal points
Blocker2510
Major4312
Minor515
Total1127

The severity-adjusted result is:

€241 ÷ 27 = €8.93 per severity point

That second number is useful when comparing test rounds. If the next round costs €241 and finds 16 issues but only 18 severity points, the raw count improved while the severity-adjusted yield got weaker.

Run the calculation in 6 steps after every test round

You do not need a research operations system to track this. A spreadsheet with one row per unique issue is enough.

  1. Write down total spend. Include session fees, incentives, recruiting costs, and the review time you want counted.
  2. Export every note from session replays and written reports. Keep timestamps when possible so engineers can replay the moment of confusion.
  3. Remove duplicates. If four testers hit the same navigation problem, keep one issue and record that it appeared in four sessions.
  4. Remove non-actionable comments. Keep opinions only when they point to a specific task failure, hesitation, misconception, or missing information.
  5. Assign severity. Use blocker, major, minor, or observation, and write one sentence explaining the rating.
  6. Calculate both metrics. Track cost per issue and cost per severity point for each testing round.

If you need help deciding what to include in a lean paid session, this breakdown of what a €29 web app testing session includes is a useful companion.

Do not reward noisy testing just because the cost per issue looks low

A low cost per issue is not always good. If a tester reports 25 vague preferences and none of them change your roadmap, your spreadsheet may look efficient while your product gains little.

Use these quality gates before counting an issue:

  • Task relevance: Did it happen during the scenario you asked the tester to complete?
  • Evidence: Can you point to a session replay timestamp, quote, or written finding?
  • Actionability: Could a designer, founder, or engineer make a specific change from it?
  • User impact: Would fixing it improve activation, comprehension, conversion, trust, or support load?
  • Repeat signal: Did more than one tester hit it, or is the single occurrence severe enough to matter?

This is why real session replays are valuable. You can see whether a “confusing button” caused a two-second pause or a complete task failure.

Compare test rounds by yield, not just by session price

Session price matters, especially for bootstrapped teams, but cheap testing that produces weak findings is expensive in disguise. Compare testing options by the number and severity of issues they uncover per euro spent.

Testing approachTypical strengthRisk to watchBest metric
Founder-led reviewFast and freeYou already know how the product is supposed to workIssues found per hour
Free user feedbackGood for broad sentimentLow control over task quality and timingActionable issues per response
Paid human testingSpecific tasks, session replays, written findingsPoor briefs can create shallow feedbackCost per severity point
Customer interviewsStrong for motivation and buying contextLess reliable for observing actual UI frictionValidated decisions per interview

If you are choosing between unpaid feedback and paid sessions, read this comparison of paid user testing vs free user feedback. The right choice depends on whether you need quick opinions, observed task behavior, or evidence before charging.

Improve the metric by writing a tighter tester brief

The easiest way to lower cost per useful issue is not always buying cheaper sessions. Often, it is giving testers a sharper task.

A weak brief says, “Look around our app and tell us what you think.” A stronger brief says, “You run a five-person agency. Sign up, create your first client workspace, invite one teammate, and stop when you know what plan you would choose.”

The second brief creates measurable moments. You can see where the tester hesitates, what they misunderstand, and whether the onboarding flow supports the job you care about.

For browser-based apps, SaaS products, marketing sites, and onboarding flows, TestTorch lets founders submit a URL and a specific scenario or brief. Sessions are performed by vetted testers who complete a screening session before accessing paid tests, and founders receive full session replays plus written findings.

When the cost per issue is high, ask these 4 questions before cutting testing

A high cost per issue does not always mean testing failed. Sometimes it means the product area is already polished, the brief was too narrow, or the tested flow was not risky enough.

  1. Did we test a high-risk flow? Pricing, signup, activation, checkout, and invite flows usually produce more valuable findings than static settings pages.
  2. Was the scenario realistic? If testers cannot imagine the job, they may browse instead of behaving like target users.
  3. Did we overcount internal review time? A first round may take 4 hours to analyze, while later rounds take 90 minutes because your tagging system improves.
  4. Did we find fewer issues because the previous fixes worked? Rising cost per issue can be a healthy sign if key blockers disappeared.

Track the trend across rounds. If round one costs €18 per severity point, round two costs €25, and activation problems are shrinking, you may be moving from discovery into refinement.

A practical benchmark for deciding whether to run another round

There is no universal “good” cost per usability issue found because the value of an issue depends on your product stage and traffic. A €60 issue that prevents trial signup can be cheap; a €5 issue about footer spacing can be noise.

Use this decision rule instead: run another round when the last round found at least one blocker or several major issues in a business-critical flow. Pause broad testing when recent sessions mostly produce minor polish items, then retest after you change the product or move to a new risky flow.

For early teams planning spend, this bootstrapped guide to human app testing cost gives a broader budgeting frame. Your metric should answer one practical question: “If we spend the same amount again, are we likely to find problems worth fixing before users hit them?”

See your own app through fresh eyes.

Post a session and get a recorded walk-through with written findings.