Insights

How to measure AI ROI: from task improvements to operational outcomes

A practical method for baselining AI value, measuring workflow impact and avoiding inflated business cases based on theoretical time savings.
AI strategy 7 min read

AI business cases often begin with a simple calculation: minutes saved per task multiplied by task volume and salary cost. It is a useful starting point. It is not a reliable return-on-investment figure.

Time saved may be absorbed by extra checking. Staff may not adopt the tool. Demand may rise to fill the capacity. Quality may improve without reducing cost. A faster task may simply move work to the next bottleneck.

A credible AI ROI case connects model performance to workflow change, then connects that change to an outcome the business can observe.

Use a four-layer measurement chain

Measure AI at four related levels. Each answers a different question.

1. Output quality

Can the system produce a sufficiently good output for the intended task?

Measures might include factual accuracy, groundedness, completeness, classification precision, policy compliance or the rate at which a reviewer accepts the output without correction.

These metrics are necessary for release decisions, but they do not show whether the business benefits.

2. Task performance

Does the tool improve the work performed by a user or team?

Measure handling time, preparation time, first-pass completion, correction effort, throughput and user success. Compare the AI-assisted task with the current method under similar conditions.

3. Workflow outcome

Does the wider process improve?

A 30 per cent reduction in document review time is of limited value if cases continue to wait in another queue. Measure end-to-end cycle time, backlog, rework, escalation, customer response time and service-level performance where relevant.

4. Economic outcome

Does the workflow change create usable financial or strategic value?

That may appear as avoidable external spend, capacity for additional volume, reduced leakage, faster revenue, improved retention, reduced remediation or lower risk exposure. Some benefits can be converted to cash quickly. Others create capacity or quality rather than an immediate cost reduction.

The UK Government guidance on impact evaluation for AI interventions recommends considering process, impact and value for money. This layered approach follows the same logic: implementation evidence and output quality are not substitutes for proof of impact.

Establish the baseline before the pilot

A baseline is the best available picture of current performance. Collect it before users know a new tool is being judged, where practical, and use the same definitions during the evaluation.

For a complaint-summary workflow, a baseline might include:

  • Median and 90th-percentile preparation time.
  • Cases completed per employee per week.
  • Proportion returned for missing evidence.
  • Quality-assurance score.
  • Time from receipt to first substantive response.
  • Staff confidence in finding the relevant facts.

Use a distribution, not just an average. Difficult cases may respond differently to AI assistance, and an apparently small tail can carry disproportionate risk and cost.

Document how the baseline was collected, which case types were included and whether demand varies by day, product or team. Otherwise a seasonal change can be mistaken for an AI effect.

Define the counterfactual

The central impact question is not “Did performance improve?” It is “What would have happened without the intervention?”

The strongest practical design depends on the workflow:

  • Randomised comparison: eligible cases or users are assigned to AI-assisted and current-process groups.
  • Phased rollout: comparable teams adopt the tool at different times.
  • Matched comparison: similar cases are compared using factors such as complexity, product and channel.
  • Before-and-after: performance is compared with the baseline, with explicit acknowledgement of other changes.

Small businesses may not have the volume for a formal experiment. They can still strengthen the evidence by using a defined sample, consistent measures and a comparison group rather than relying on impressions.

Separate gross benefit from realised value

Start with the measurable change, then apply the operating conditions that determine whether the benefit is realised.

A simple capacity calculation is:

Gross capacity released = eligible volume x adoption rate x net time saved per case

Net time saved includes review, correction and exception handling. It should not use the model’s generation time alone.

Then ask how the capacity will be used. There are several legitimate answers:

  • Absorb growing demand without equivalent hiring.
  • Reduce overtime or contractor spend.
  • Reallocate time to higher-value work.
  • Improve service levels or reduce backlog.
  • Increase the amount of quality checking.

Only count a direct cost saving if the cost will genuinely change. Capacity and cost reduction are different benefits and should be presented separately.

Build a complete cost model

AI cost is more than model usage. Include:

  • Discovery, workflow design and data preparation.
  • Integration and security work.
  • Licences, hosting and model usage.
  • Evaluation design and human review.
  • Change management and training.
  • Monitoring, support and incident response.
  • Vendor management and future migration.
  • The expected cost of errors and exceptions.

Use a realistic planning horizon. A low-cost prototype can require substantial work to meet production standards. Conversely, a reusable evaluation service or integration can support several later use cases, so avoid loading every shared cost onto the first workflow without explanation.

Measure quality and risk as first-class outcomes

An AI workflow may be valuable because it makes work more consistent, not because it makes it cheaper. Quality benefits can often be quantified.

Examples include fewer missing fields, higher audit scores, better policy adherence, lower rework, fewer inappropriate escalations or earlier identification of vulnerable customers.

Risk reduction is harder to value because avoided events are not directly observed. Use leading indicators and scenarios instead of pretending to know an exact cash value. Track override rates, high-severity errors, control failures and near misses. For major risks, show a plausible range rather than one confident number.

The NIST AI Risk Management Framework treats measurement and management as continuing activities. That is a useful model for ROI too. Value can deteriorate when data, user behaviour, model versions or business processes change.

Add adoption to the business case

An accurate tool with poor adoption has little operational value. Track:

  • Eligible cases where the tool was used.
  • Users active by team and role.
  • Outputs accepted, edited or rejected.
  • Reasons for non-use and workarounds.
  • User confidence and training needs.
  • Whether the tool creates additional steps.

Do not target 100 per cent adoption automatically. Some cases may be outside the approved scope. The meaningful measure is appropriate use on eligible work.

Use a value scorecard

Keep one scorecard with leading and lagging indicators.

Quality: task success, acceptance without correction, severe error rate.

Efficiency: net time per eligible case, throughput, queue age.

Outcome: cycle time, rework, service level, customer or employee outcome.

Adoption: eligible use rate, active users, override and rejection reasons.

Cost: build and run cost, human review cost, cost per successful case.

Risk: control failures, incidents, complaints and exceptions outside tolerance.

Every metric needs a definition, data source, owner and review cadence. Set a target and a stop threshold before launch.

A practical worked example

Suppose a team handles 2,000 eligible cases per month. The baseline preparation time is 18 minutes. In a controlled comparison, AI assistance reduces median preparation time to 12 minutes, including checking and corrections. Appropriate adoption is 70 per cent.

The gross monthly capacity released is 2,000 x 70 per cent x 6 minutes, or 140 hours.

That is not automatically 140 hours of salary saving. The team may use the capacity to reduce a backlog, respond sooner or avoid additional contractor support. The business case should state which outcome is expected and how it will be verified.

It should also show quality. If acceptance without material correction is below the release threshold, or severe omissions increase, the six-minute gain is not a success.

Make the investment decision explicit

At each gate, choose one of three outcomes:

Proceed: evidence meets quality, risk and value thresholds, and an owner has committed to the operating change.

Refine: the use case remains valuable but needs a narrower scope, better data, different controls or workflow changes.

Stop: the likely realised value does not justify the cost and risk, even if the prototype is technically impressive.

Stopping is a valid result. It prevents a weak use case from consuming more investment and improves the evidence behind the next choice.

Treat ROI as an operating measure

Recalculate the value case after launch using real adoption, real volume, actual review effort and current running cost. Review it when the model, workflow or business conditions change.

The purpose of AI ROI is not to create a large number for approval. It is to help the organisation decide where to invest, what to change and when to stop.

For help defining baselines, evaluation measures and a defensible value case, see our AI strategy and delivery services or talk to Sorsana.

  • AI ROI
  • Business case
  • Measurement