Build the Measurement Two Weeks Before the Pilot, or You'll Never Have a Number
Build the Measurement Two Weeks Before the Pilot, or You'll Never Have a Number
More than half of the industry can't evidence what its AI deployments are worth. The reason is almost never the technology — it's that nobody wrote down what "before" looked like.
There is a board meeting that happens in a lot of organisations about nine months after an AI tool goes in. Someone asks what it's delivered. The answer that comes back is a story: the team says it's much faster, people like it, the month-end feels smoother. All of that may be perfectly true, and none of it is a number.
The finance leader in the room is usually the one who ends up wearing that. And the honest position is that the measurement problem was created long before the meeting — on the day the pilot started without a baseline.
How widespread this is
The Cambridge Centre for Alternative Finance published its 2026 Global AI in Financial Services Report at the end of April, drawing on 628 respondent organisations across 151 jurisdictions — fintechs, incumbent financial institutions, AI vendors, and 130 central banks and regulators.
Its finding on measurement is blunt. Fifty-five per cent of industry respondents and 63 per cent of surveyed regulators find it difficult to measure the value of AI deployment, rising to 76 per cent among large financial institutions. Alongside that, only 40 per cent report increased profitability from AI while 43 per cent report no change at all.
Two things about that data are worth stating plainly. It surveys financial services rather than the not-for-profit, disability and care sectors most readers here work in, so the percentages aren't yours. And it's from April, not last week. What travels is the shape of the finding, not the figure — and the shape is that difficulty rises with organisational size, which suggests resourcing alone doesn't explain it. Bigger organisations have more analysts, better systems and more governance. They also have more process between the change and the outcome, and that's what breaks the measurement.
|
55% → 76%
Industry respondents finding it difficult to measure the value of AI deployment, rising to 76% among large financial institutions. Cambridge CCAF, April 2026.
|
40% / 43%
Report increased profitability from AI, against those reporting no change at all — productivity effects are felt well ahead of enterprise value being evidenced.
|
Why the number doesn't exist
The usual explanation is that AI benefits are soft and hard to quantify. I don't think that's right, or at least it's not the binding constraint.
The binding constraint is that nobody captured the "before". A pilot begins because someone is enthusiastic, and enthusiasm doesn't generate a two-week measurement exercise about how long things currently take. The baseline is boring; the pilot is interesting. So the tool goes in, the process changes, and three months later the only honest comparison available is between today's measured state and a remembered one.
That failure is structural rather than analytical. It cannot be fixed afterwards by better analysis, because the data simply isn't there. It can only be fixed by doing something unglamorous first — which is the entire argument of this post.
The four numbers to capture before anything is switched on
Two weeks of ordinary operation, before the tool exists. Four measures, all of which can be collected by hand if necessary. Scale the effort to the team: in a finance function of three or four people this is a tally sheet and a few diary notes, not a project.
Cycle time. Elapsed hours from a task arriving to it being finished and accepted — not effort, elapsed time. Effort measures how busy someone was; cycle time measures what the organisation actually experiences. For a claims batch, an acquittal, or a month-end task, this is usually the number a funder or a board cares about.
Rework rate. The proportion of outputs that come back for correction after being considered complete. This is the measure that most often moves in the wrong direction after an AI tool is introduced, and the one nobody notices because it was never counted before.
Review time. How long a second person spends checking the work. Worth isolating because AI frequently shifts effort here rather than removing it — a draft produced in two minutes and reviewed for forty is not obviously better than one produced in twenty and reviewed for ten.
Error escape rate. How many errors reach the point of no return — a lodged return, a submitted claim, a distributed report. In funded sectors this is the one with real consequences attached, and it should be tracked separately from rework because they behave differently.
Collect these for two weeks, write them on one page, and date it. That page is worth more than any vendor's benefit case, because it's about your organisation.
Why cost per task is the wrong headline
Almost every vendor pitch converts to cost per task or per document, and it's an appealing metric because it's easy to compute and always looks good.
It's the wrong headline for two reasons. It ignores the review layer, which is where a genuine share of the cost lands and where poor-quality output is most expensive. And it treats every task as equivalent, which in a funded organisation they emphatically are not — a mis-stated grant acquittal and a mis-coded stationery invoice are not two units of the same thing.
The more useful framing is cost per accepted output: total cost including review, divided by outputs that made it through without correction. That number is harder to game, and it's the one that moves when quality slips.
Write the kill criterion before you start
The one discipline that most reliably separates a pilot from an open-ended subscription is deciding, in advance and in writing, what result would cause you to stop.
Something like: if after twelve weeks cycle time hasn't improved by at least a quarter, or if rework rises at all, we discontinue. Specific enough to be uncomfortable. Agreed by whoever sponsors the pilot, not by the person running it.
This matters because pilots almost never fail — they get extended. Once a tool is embedded in a workflow, the cost of removing it starts to feel higher than the cost of keeping it, and the assessment quietly changes from "is this working" to "what would we do instead". A kill criterion set at the start is the only thing that reliably interrupts that. And a pilot you were genuinely willing to stop is the only kind that tells you anything.
Where AI helps with its own measurement
Some of this is tedious enough that it doesn't get done, which is a legitimate use for the technology itself. Extracting timestamps from a workflow system to derive cycle time, classifying corrections into categories to produce a rework rate, summarising a fortnight of activity logs into a baseline table — all structured, high-volume, low-judgement work.
What stays human is deciding what counts as an error, what counts as accepted, and what threshold would justify stopping. Those are the definitions the whole exercise rests on, and they can't be delegated to the thing being assessed.
Could you evidence what your AI tools have delivered?
If the answer is a story rather than a number, the baseline is missing — and the next pilot is the last easy chance to build one. PFL provides senior-level outsourced finance, management reporting, and AI automation for Australian NFP, NDIS, and SME organisations.
Talk to PFL →Report finds uneven AI adoption in financial services — Cambridge Judge Business School, 28 April 2026
The 2026 Global AI in Financial Services Report: Adoption, Impact and Risks — Cambridge Centre for Alternative Finance
Voluntary AI Safety Standard — Australian Government Department of Industry, Science and Resources
Guidance on privacy and the use of commercially available AI products — Office of the Australian Information Commissioner
Comments
Post a Comment