Define value before measuring automation

Name the outcome in operational or financial terms: cases resolved correctly, time to an approved deliverable, avoided loss, additional contribution or a learning result. State who benefits and over what period. A count of generated messages does not show that the messages helped.

For hard-to-price outcomes, report a balanced scorecard rather than inventing a monetary value. Quality, cycle time, accessibility and risk can be decision-relevant without being forced into one number.

Use a full-cost ROI formula

Total cost can include usage, tools, infrastructure, integration, evaluation, monitoring, review, corrections, training and failure. Separate one-time implementation cost from recurring operating cost so the payback story remains visible.

Match evidence to the claim

Claims and useful evidence
ClaimEvidence approachCaution
Work became fasterComparable cycle-time baselineCheck whether quality or backlog changed
Cost fellCost per accepted task before and afterInclude review and retries
Outcome improvedControlled test or credible comparisonCorrelation alone may reflect other changes
Risk declinedDefined incident and severe-failure measuresRare events need longer observation

Randomised holdouts can support causal claims when feasible. When they are not, disclose the comparison's limits and avoid attributing every change to AI.

Evidence & context: Google Research

Build a small decision dashboard

  • Volume attempted and accepted.
  • Quality and severe-failure rate by task category.
  • Cost and time per accepted task.
  • Retry, escalation and human-review rates.
  • Business or learner outcome tied to the original objective.
  • A baseline, target, owner and review date.

Use the dashboard to stop as well as scale. If value does not exceed total cost at the required quality threshold, redesign the task or retire the workflow.

Sources & further reading

  1. Methods for Measuring Brand Lift of Online Ads

    Google Research. Original research using randomised experiments to estimate advertising effects; no universal lift or ROI benchmark is inferred.

  2. Generative Artificial Intelligence Profile (NIST AI 600-1)

    NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.

  3. Evaluate a model router

    Microsoft Learn. Official guidance for evaluating routing across representative workloads using quality, cost, latency and policy criteria. It is not evidence that routing always improves results.

Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.

A correction, a counterexample or an experience worth sharing?

Join the conversation ↗