Measuring it so the numbers survive
A CFO will ask what you got for the money. This lesson is how to have an answer that holds up, and how to avoid the three ways these numbers go wrong — all of which produce a result that is too good and falls apart under questioning.
Three numbers, not one
Time. Minutes per unit of work, before and after. The easiest to measure and the most often quoted alone, which is how a deployment that halved the time and doubled the error rate gets reported as a success.
Quality. Error rate, rework rate, complaint rate, QA sample scores — whatever your operation already uses. Use the existing measure rather than inventing one, because the existing one has a history to compare against.
Volume. Throughput per person per week. This is the one that reveals whether saved time became more work done or simply evaporated, which is a real and common outcome and not necessarily a bad one — but you should know which happened.
Report all three together, always. A change in one is not interpretable without the other two.
The three ways these numbers lie
You measured the enthusiasts. Volunteers for an AI pilot are people who like AI, are better at it, and want it to work. Their results overstate what the team will get by a wide margin. Include people who did not volunteer, and report the range as well as the average — a mean of 35% time saved that runs from 5% to 70% is a different fact from one that runs from 30% to 40%.
You did not count the checking. The output arrives in one minute. Verifying it takes eight. If your metric starts when the person opens the tool and ends when the draft appears, you have measured the wrong interval. Measure end to end, from the work arriving to the work being finished and trusted.
You measured the novelty. Weeks one and two of any new tool are unrepresentative — inflated by enthusiasm, deflated by unfamiliarity, and never a steady state. Report from week three onwards, and re-measure at three months, which is when the honest number appears.
The costs that never reach a spreadsheet
Review time, as above. It is the single most under-counted quantity in this entire field.
Vigilance decay. This is the subtle one and it is worth understanding. When a system is right 92% of the time, human checkers get worse at catching the 8% — because attention degrades when errors are rare, which is a well-studied effect in every domain where people monitor mostly-reliable automation. Your error rate at month six may be higher than at month one with nothing about the system having changed. Design for it: rotate checkers, sample deliberately, and re-measure quality on a schedule rather than assuming it holds.
Process debt. The exceptions the tool cannot handle accumulate somewhere, usually with one experienced person who becomes a bottleneck nobody planned.
What to actually report
One page. Baseline, result, and the three numbers with their ranges. What you changed in the process, not just what you bought. What broke and what you did about it. What it cost, including internal time. And a plain recommendation with the confidence you actually have.
Kavita's page said: review time 18 → 11.5 minutes (range 8–16), QA errors flat at 2.1%, throughput up 22%, 15% of files now flagged for source verification, licence cost ₹X, internal time 60 hours over eight weeks, recommend proceeding for this task only.
That last clause — for this task only — is what made the page credible. A report recommending AI everywhere reads as advocacy. One recommending it for a specific task after measuring it reads as work, and it is what got her the budget for the next one.
The question worth asking at six months
Would we go back?
Ask the team, not the sponsor, and ask it as a real question. It cuts through metric ambiguity better than anything else, because people who would not go back are telling you the tool is genuinely load-bearing, and people who shrug are telling you it is not — regardless of what the time numbers said.
Do this today: find out whether you have a baseline for your top task. If you do not, start measuring it this week. Everything else waits on that.