A company can have widespread AI use and little idea whether it is getting a return. Employees report saving time, subscription numbers rise and demonstrations improve. Yet the customer still waits, the backlog remains and the operating budget absorbs another expense.
The problem begins with what management counts. Access is measurable immediately. A change in operating performance requires a baseline, a defined process and someone accountable for the result.
I would separate three claims that are too often treated as interchangeable: an individual task became faster, the whole process improved, and the business obtained a valuable result. Each needs its own evidence.
The Workflow
A task-level gain has to survive the rest of the workflow
Consider a hypothetical proposal process. AI reduces the effort required to draft a document, but pricing exceptions still wait for a manager and contractual language still moves through the same review queue.
The drafting team has gained capacity. The customer may receive nothing sooner. If the additional drafts increase the review backlog, the downstream team can become worse off.
A useful business case follows the proposal through acceptance, including revisions and approvals. It asks whether the company can handle more qualified opportunities, respond earlier or spend less on each completed proposal without losing commercial control.
That is a different question from asking how many minutes a user believes the tool saved.
The Evidence
The evidence supports selective deployment
Research gives good reasons to take the opportunity seriously and equally good reasons to avoid uniform promises. In the revised version of Generative AI at Work, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,172 customer-support agents. They reported an average 15% increase in issues resolved per hour, with substantially different effects across workers.
Less experienced and lower-skilled workers improved speed and quality. The most experienced and highest-skilled workers showed small speed gains and small quality declines.
The useful inference is about deployment design. A capability that helps a newer employee handle an unfamiliar problem may add little to an expert’s established method. Training, review and expected benefit should reflect that difference.
The study establishes a result in a particular operating setting. It does not establish a 15% saving for every department that buys an assistant. Your workflow, customers, data and quality requirements still need to be tested.
The Advantage
Operating improvements create commercial choices
Lower cost per completed case can create room to reduce prices, invest in service, absorb more volume or retain margin. Shorter cycle time can let a company respond while an opportunity is still open. More consistent quality can reduce rework and the senior attention spent recovering failures.
Those choices explain why AI can become a competitiveness issue. A rival does not need a more impressive model if it can deliver the same customer promise with a better operating process.
The benefits can also accumulate. Recovered capacity supports more work; faster feedback helps refine the next cycle. But this is a mechanism to evaluate, not a reason to assume every competitor is already achieving it.
THE THROUGH-LINE
The relevant comparison is specific: which part of your cost, speed or reliability would become a disadvantage if another firm improved the equivalent workflow?
The Baseline
Set the measure before choosing the system
Start with a process whose current weakness is visible. Record its volume, elapsed time, labor effort, error rate and exception handling. Choose a primary result and the quality or risk limits that must hold alongside it.
In customer support, faster first responses can conceal more reopened cases. In finance, faster extraction can create more reconciliation work. In sales, more proposals can conceal weaker qualification or greater pressure on margin.
The measure should therefore include the work the automation generates elsewhere. Count review, correction, integration support and the continuing maintenance of the system. Include the process owner’s time when that becomes a recurring requirement.
Where practical, compare similar cases over comparable periods or introduce the change in stages. A quieter month, a change in case mix or a new staffing pattern can otherwise be mistaken for an AI effect.
The resulting estimate may be a range. That is more useful than a precise saving built from assumptions nobody has tested.
The Capacity
Recovered time needs an operating decision
Ten minutes saved in many scattered moments does not automatically remove ten minutes from payroll. It may make the job less pressured, improve responsiveness or create capacity that cannot be scheduled efficiently.
Management needs to decide how the recovered time becomes useful. Does the team take more cases? Reduce overtime? Improve the depth of review? Stop a duplicated task? Each route has a different economic effect.
Treat an avoided hire, a reduction in paid hours and additional capacity as separate outcomes. They can all matter. They cannot all be counted as the same saving.
This is why the workflow owner belongs in the investment decision from the start. The technical team can enable a different way of working. It cannot alone decide how another department uses the capacity.
The Timing
Waiting can be a disciplined choice
AI may be a poor investment for an infrequent task, unstable process or output whose correctness is expensive to establish. Improving data quality or simplifying a rule may be the better next move.
A decision to wait should name the missing condition and a trigger for reconsideration. That could be an acceptable error rate, sufficient volume, a viable integration or a control the current product lacks.
Leaving the question unowned is different. It preserves the existing cost without requiring anyone to explain why that cost remains acceptable.
The Call
Use the first 90 days to reach a defensible decision
For a suitably bounded workflow, a 90-day evaluation can produce something more useful than a portfolio of demonstrations. Establish the baseline, appoint the owner, test representative work and assess the full operating consequence.
At the review, management should be able to explain whether the target result improved, what it cost to obtain, which failures appeared and whether to expand, revise or stop. If the relevant business cycle is longer, say so and choose an evaluation period that can actually produce evidence.
The next investment should follow that decision. More licenses and more pilots are justified when they extend a demonstrated benefit or test a specific uncertainty. Counting them as success before either has happened leaves the business paying for activity it cannot value.