Baseline It First
How to measure AI automation ROI: only 39% of companies can tie any profit to AI. The fix is writing down three numbers the week before you start.
Hören Sie auf zu konfigurieren. Fangen Sie an zu bauen.
SaaS-Builder-Vorlagen mit KI-Orchestrierung.
Problem: Your automation went live, the team says it is saving hours, and your CFO asks a simple question you cannot answer: hours saved compared to what? Nobody wrote down how long the work took before. So the return on investment is now a debate instead of a number.
Quick Win: A baseline is the set of numbers you write down about a process before you change anything. It is the cheapest thing you can do to make an AI project provable, and almost nobody does it. Capture three numbers for two to four normal weeks before you start: how long one piece of work takes from request to done, how many finished pieces one person produces per week, and how often finished work has to be redone. That is the whole method. If you did not measure the thing before you changed it, you do not have a return on investment, you have an anecdote.
Why Nobody Can Prove Their AI Paid For Itself
The numbers come from the biggest survey in the field. McKinsey's State of AI, published November 2025, covered 1,993 respondents across 105 countries. 88% of organizations now use AI in at least one part of the business. But only 39% report any operating profit they can attribute to AI, and most of those put it under 5%. About 6% attribute more than 5% of operating profit to AI (McKinsey, summarized by CX Today).
Read that carefully. It does not say 61% of AI projects failed. It says 61% cannot connect AI to profit at all. A project can be working beautifully and still be unprovable, because nobody captured the before.
Gartner sees the same pattern from the finance side. In a March 2026 survey of 204 finance leaders, 66% of teams that had adopted AI named efficiency and productivity as a top benefit, while 63% said their rollouts ran slower than expected in 2025. Gartner's warning to CFOs names the confusion directly: "They must not mistake activity for impact. Counts of pilots, tools rolled out or use cases in production show that finance is moving, but they do not prove that AI is delivering the value boards now expect" (Gartner).
Productivity is the easiest thing to claim and the hardest thing to prove. That is exactly why so much money goes there and so little comes back.
The Week Before: What To Write Down And Where To Get It
Here is the part that takes four hours and saves you a year of arguing. Pick the one process you are changing and capture it cold, for two to four normal weeks. Not your busiest week. Not the week of the holiday. Normal.
| What to capture | In plain English | Where to get it | How long |
|---|---|---|---|
| Cycle time | How long one piece of work takes from request to finished | Timestamps in your ticket system, the software you use to track customers and deals, or email threads | 2 to 4 weeks, or the last 30 items |
| Finished work per person per week | How many completed proposals, quotes, invoices, or tickets one person gets out the door | Count finished items, divide by the people doing that work | Same window |
| Rework rate | How often finished work comes back to be fixed or redone | Version counts, "revised" file names, reopened tickets, refunds, rejections | Same window |
| Full cost per item | Salary plus overhead for the people touching it, divided by the number of items | Payroll, plus whatever your finance team adds on top for overhead | One-off calculation |
| Volume | How many of these happen per week | Any of the above sources | Same window |
Two rules make this stick.
Pull from records that already existed. Timestamps, invoice dates, ticket open and close times, file version history. Nobody was gaming those numbers, because nobody knew they would be used this way. The moment you ask people to self-report their own hours, the baseline is contaminated. The next section shows how badly.
Get finance to sign it before day one. A number your CFO agreed to in week zero cannot be argued away in month six. Same discipline as a 30-day pilot with a go or no-go decision agreed in advance: agree the measuring stick before anyone has a reason to prefer one answer.
Three Numbers That Survive A Board Room
Most AI reporting drowns a board in metrics that sound like effort. Prompts run. Tools deployed. Adoption rate. None of those are outcomes. Three numbers survive the room, because each is already on a line the board understands.
1. Cycle time. How long the work takes from request to done. The cleanest metric in the set, because machines capture it rather than people, and because a board already knows what a shorter sales cycle or a faster month-end close is worth. Ten-day close becomes six-day close: that is a fact with a date attached.
2. Finished work per person per week. Not activity, finished units. Proposals sent. Invoices issued. Tickets resolved. This is the number that answers "do we need to hire," and hiring decisions are the only place most productivity claims ever touch the profit line.
3. Rework rate. The percentage of finished work that has to be redone. This is the honest one, because automation that doubles output while tripling the fix-it pile is not a win, it is a hidden cost moved onto whoever checks the work. The American Society for Quality's widely cited estimate puts total quality-related costs at 15 to 20% of sales revenue at many manufacturers (Fabrico). That is not small money, and checking machine output is the same category of cost.
Rework rate also catches the failure the demo hides. In the Upwork Research Institute's survey of 2,500 executives, employees, and freelancers across the US, UK, Australia, and Canada, 39% of employees using AI said they were spending more time reviewing and moderating AI output (Upwork). Untracked, that review time is invisible on your scoreboard and very visible in your team's week. Same cost we break down in what a 1,000-agent workflow costs to run: the review is the expensive part, not the software.
Why "Hours Saved" Is The Wrong Unit
Hours saved is the metric every AI pitch leads with and no board accepts. Two reasons, both fatal.
People are bad at estimating their own time savings. METR ran a controlled study with 16 experienced open-source developers across 246 real tasks, in codebases they had worked in for an average of five years. Each task was randomly assigned to allow or forbid AI tools. Afterwards, the developers estimated AI had made them about 20% faster. The clock showed they were 19% slower (METR).
A 39-point gap between what skilled professionals believed about their own work and what actually happened. The tools in that study were the ones available in early 2025, and METR itself now treats the result as a snapshot of that moment rather than a verdict on today's tools. Take the narrow lesson, which has not aged: an hours-saved figure built on impressions carries an error bar wide enough to swallow the business case.
A saved hour is not money. It is money only if someone decides what it becomes. The same Upwork study found 96% of executives expected AI to raise productivity, while 77% of employees using AI said the tools had added to their workload, and 47% said they did not know how to get the gains they were supposed to be getting (Forbes).
The finance version is simple. Cut a ten-day month-end close to six days. If the four recovered days get absorbed into meetings and email, the profit line does not move. The productivity gain is real. The business impact is zero.
Turning A Saved Hour Into A Real Decision
Saved capacity becomes money in exactly three ways. Pick one before you start and write it into the project brief.
| Saved capacity becomes | What has to happen | The number that proves it |
|---|---|---|
| Revenue | Freed hours go into work that bills or sells, and you can point to which work | Revenue or billable hours per person, same team, before and after |
| A hire you do not make | Volume grows and headcount does not, or a planned role gets cancelled | Volume per person, plus the cancelled role |
| A cost you stop paying | A tool, a contractor, or an outsourced step is switched off | The invoice that stops arriving |
If your project does not map to one of those rows, it does not have a return. It has a nicer working day, which is genuinely worth something and is not a number you take to a board.
Name the row on day one. "This exists so we can handle 40% more proposals without adding a second proposal writer" is a testable claim. "This will save the team time" is not. That is the difference between recurring internal work produced automatically and a tool nobody can defend at budget time.
The One-Page Before And After
Everything above collapses into one page that fits on a single screen. This is the format we hand back. The numbers below are illustrative placeholders to show the layout, not client data.
Process: Proposal production Baseline window: 4 weeks, 12 March to 8 April, pulled from the deal-tracking system Signed off by: Finance, 9 April
| Before | After (4 weeks) | Change | |
|---|---|---|---|
| Cycle time, request to sent | 6.2 days | ||
| Proposals finished per writer per week | 4.1 | ||
| Rework rate (sent then revised) | 31% | ||
| Full cost per proposal | $410 | ||
| Volume per week | 21 |
The claim being tested: volume rises above 30 per week with no additional writer. The decision it feeds: cancel the second proposal-writer hire budgeted for Q4.
That page beats every dashboard you will be sold, because it commits to an answer before anyone knows what the answer is. Note the empty columns. Filling them in later is what makes it evidence instead of marketing.
Honest Failure Modes
This method breaks in specific ways. Knowing them in advance is the difference between a measurement and a fight.
The process is genuinely unmeasurable. Some work has no timestamps and no countable unit. Strategy, relationship building, judgment calls. If you cannot define the unit of finished work, do not fake a baseline. Run a matched comparison instead: keep one team or one queue on the old process for four weeks and compare the two. Slower, less tidy, honest.
The window was not normal. A baseline captured during quarter-end or a holiday stretch flatters or punishes the automation for reasons unrelated to it. If you realize this afterwards, say so and start the baseline again. A quietly wrong baseline is worse than none, because it survives longer.
Something else changed at the same time. You automated proposals in the same quarter you hired two salespeople and changed pricing. Now you cannot tell which change did what. Change one thing at a time, or label the result a strong hint rather than proof.
The metric gets managed instead of the work. Once cycle time is the scoreboard, people find ways to make cycle time look good, usually by pushing work into a queue nobody measures. That is why rework rate belongs in the set. It is the counterweight that catches the shortcut.
The baseline proves the automation did not work. The one people quietly fear, and the whole point. MIT's NANDA research found roughly 95% of company AI pilots produced no measurable profit impact (Fortune on MIT NANDA). A baseline is how you find out in month two instead of year two. Killing a project early is a return. It shows up as money you did not spend.
Related Reading
- The 30-day AI pilot that ships, where the go or no-go number gets agreed before day one
- What a 1,000-agent workflow costs to run, the review and upkeep costs your baseline has to capture
- Business process mapping to find the constraint, choosing which process to measure at all
Frequently Asked Questions
What is a baseline in an AI or automation project?
The numbers describing a process before you change it: how long the work takes, how much gets finished per person, how often it has to be redone, and what it costs. Captured from records that already exist, over a normal two to four week window, signed off by finance before the project starts. Everything you claim afterwards is measured against it.
Which metrics should we report to the board?
Cycle time, finished work per person per week, and rework rate, each as a before and after with the same source and the same window. Then one line naming the decision those numbers feed: a hire not made, a cost switched off, or revenue produced with the same team. Gartner's framing is the useful one: pilots launched and tools deployed show activity, not impact (Gartner).
Is hours saved ever a useful number?
As an internal signal, yes. As proof to a board, no. Use it to spot where to look, then convert it into one of the three things that move the profit line. And never take it from self-reports alone: developers in METR's study believed AI made them 20% faster while measurement showed 19% slower (METR).
Most companies buy the automation first and go looking for proof afterwards, which is why 61% of them cannot connect AI to profit at all. We do it the other way around: measure the process cold, agree the number the work has to hit, build the smallest thing that hits it, then hand back a one-page before and after your CFO signed off on in advance. See what we build for companies →
Hören Sie auf zu konfigurieren. Fangen Sie an zu bauen.
SaaS-Builder-Vorlagen mit KI-Orchestrierung.