AI workflow ROI should be measured against accepted work, not licences issued, prompts sent or tokens consumed. Start with one operational outcome—such as a correctly routed request, a reviewed evidence pack or a resolved service case—define the quality bar, then calculate the full cost of producing outcomes that meet it. That makes retries, corrections, human review and escalations visible instead of hiding them behind an adoption number.
This matters because AI activity can grow while operational value stays flat. A team can use an assistant every day and still create more checking, rework or exception handling than the workflow saves. The practical buyer question is not “are people using AI?” It is “are we completing useful work more reliably, quickly and economically?”
What is cost per accepted outcome?
Cost per accepted outcome is the total cost of running an AI-assisted workflow divided by the number of outputs that met a defined quality bar.
The numerator should include more than the model bill:
- model and tool usage
- integration and platform costs attributable to the workflow
- failed attempts and retries
- employee time spent reviewing or correcting outputs
- time spent resolving exceptions or manually finishing work
The denominator is not every output generated. It is the number accepted for operational use: cases resolved, documents approved, records correctly updated, or another workflow-specific definition of “done”.
OpenAI’s July 2026 investment guidance recommends measuring model and tool use, attempts, completion rate, latency and human review against a predefined quality bar. For priority workflows, it advises tracking cost per accepted outcome and pairing it with a business measure such as cycle time, capacity created, revenue protected or risk avoided.
Source: OpenAI, How to manage AI investments in the agentic era, 14 July 2026.
Why licences, tokens and time saved are not enough
Licences and active users show reach. Token cost shows one input to delivery. Estimated time saved can indicate potential. None of them, alone, tells an operations leader whether the work reached the required standard.
A cheaper model may require more retries or more human correction. A highly used tool may be helping with low-value tasks. A workflow may look fast until its exceptions reappear as a manual backlog elsewhere. Measuring only the visible AI activity rewards local efficiency while missing the end-to-end result.
The market is already showing this value-readiness gap. SAP and Oxford Economics surveyed 2,600 business leaders across 13 countries, including Australia. Respondents expected average AI ROI to rise from 16% last year to 21% in 2026 and 38% in two years. Yet only 3% said their organisation was fully prepared for agentic AI. The same research found 79% had experienced rework, delays or backlogs due to low-quality AI outputs.
Source: SAP News Center, Business value of AI is spiking, driven by increased adoption and agentic expectations, 15 July 2026.
The lesson is not that AI investment lacks value. It is that optimistic portfolio-level ROI needs workflow-level evidence underneath it.
Define three quality states before calculating ROI
Before a pilot starts, agree what happens to each completed output. A useful minimum classification is:
| Quality state | Operational meaning | What to count |
|---|---|---|
| Ready to use | Meets the quality bar as delivered | Accepted outcome |
| Needs correction | Usable after another attempt or human edit | Rework time and retry cost |
| Needs escalation | A person must investigate or finish the work | Escalation volume, delay and handling cost |
These states turn vague “accuracy” discussions into operational evidence. A workflow with a 70% ready-to-use rate and a short, predictable correction path may be commercially useful. A workflow reporting 95% model accuracy could still be expensive if the remaining failures are hard to detect or costly to recover.
OpenAI’s proposed “Useful Intelligence per Dollar” scorecard uses the same distinction between ready to use, needs correction and needs escalation. It argues that the full cost of a successful task includes employee time, human review, retries and rework—not just compute.
Source: OpenAI, A scorecard for the AI age, 17 July 2026.
Build an AI workflow ROI scorecard operators can use
A useful scorecard should fit on one page and be reviewed by the person who owns the workflow. Start with six measures:
- Accepted outcomes: the volume of work meeting the agreed quality bar.
- Ready-to-use rate: accepted without correction as a percentage of completed outputs.
- Correction and escalation rate: where residual work is being pushed onto people.
- Full workflow cost: model, tools, retries, review and exception handling.
- Cost per accepted outcome: full workflow cost divided by accepted outcomes.
- Business effect: cycle time, service level, capacity, revenue, error reduction or another relevant operational result.
Add two guardrails beside the economics: incidents and control breaches. A lower unit cost is not a win if the workflow accesses prohibited data, takes an unauthorised action or makes failures harder to detect.
The scorecard should also state the baseline. Compare the AI-assisted path with the current manual or rules-based path using the same outcome definition. If the old process measures completed evidence packs after review, the new process should not claim every draft as a success.
Measure the whole workflow, not an isolated AI step
Real work crosses systems and handoffs. AI may classify an incoming request, retrieve context, prepare a draft, route it for approval and update a record. Optimising only the drafting step can move cost into review, reconciliation or customer recovery.
Cars24 provides a current example of workflow-level implementation. Its customer agents support a journey across qualification, recommendations, bookings, finance details and follow-up. The company reports more than one million customer-conversation minutes handled by AI each month. Internally, it has deployed ChatGPT Enterprise and Codex to about 600 central employees and reports 85%–90% daily active usage.
Source: OpenAI, How Cars24 scales conversations and builds faster with OpenAI, 16 July 2026.
Those adoption numbers show scale, but the more useful implementation pattern is the connection to repeatable workflows. For another organisation, the right outcome may be a completed intake, an approved purchase request or an exception returned to the correct owner—not a conversation minute or an active user.
A practical 30-day measurement plan
You do not need a finance transformation to test this method. Use one workflow and one accountable owner.
Week 1: define the outcome. Write down what “accepted” means, who decides, which cases are in scope and which must escalate. Capture the current volume, cycle time, handling effort and error or rework rate.
Week 2: instrument the workflow. Log each completed output, its quality state, attempts, review time, escalation reason and direct AI/tool cost. Keep the data simple enough that the operating owner will maintain it.
Week 3: run representative work. Include normal cases and known exceptions. Do not remove difficult cases merely to improve the dashboard. Record where poor source data, unclear instructions or missing authority causes failure.
Week 4: review the economics and controls. Calculate cost per accepted outcome, compare it with the baseline and inspect the correction and escalation queue. Decide whether to scale, redesign, constrain or stop the workflow.
Set the decision thresholds before the review. For example: a minimum ready-to-use rate, a maximum correction time, zero tolerance for defined control breaches and a target improvement in cycle time or unit cost. This keeps the pilot from becoming an open-ended demonstration.
Where Rettare starts
Rettare starts with the workflow and the quality bar, then designs the automation, approvals, logging, exception paths and ownership around it. The point is not to prove that a model can produce an output. It is to establish whether the organisation can accept that output reliably and improve the economics of real work.
If your AI dashboard reports activity but cannot show accepted outcomes, corrections, escalations and full delivery cost, fix the measurement model before adding more workflows.
References
- OpenAI, How to manage AI investments in the agentic era, 14 July 2026.
- OpenAI, A scorecard for the AI age, 17 July 2026.
- SAP News Center, Business value of AI is spiking, driven by increased adoption and agentic expectations, 15 July 2026.
- OpenAI, How Cars24 scales conversations and builds faster with OpenAI, 16 July 2026.