Nearly Four Months of AI Inside Excel: What It Catches in a Model, and What It Can't

A grid of cells where structural breaks are highlighted clearly while one cell containing a wrong-but-well-formed value stays unmarked

Nearly Four Months of AI Inside Excel: What It Catches in a Model, and What It Can't

Long enough to stop speculating. The strengths and the blind spots are both consistent — and they divide along a line that tells you exactly how to use the thing.

In its announcement of 5 May, Anthropic confirmed the Claude add-ins for Microsoft 365 as generally available across Excel, Word and PowerPoint, alongside ten prebuilt agent templates aimed at finance work — model building, general ledger reconciliation, month-end close, statement auditing and others. Microsoft's own Copilot has been in the same surface considerably longer, and now lets you choose the underlying model. That's close to four months of live use in the one case and well over a year in the other.

This isn't a news post, and there's no August development to report. It's the assessment that only becomes possible after the launch coverage has faded: what these tools actually do to a workbook when a finance team points one at a model it depends on.

The line the capability falls along

The pattern is consistent enough to state as a rule. These tools are strong at anything that can be determined from the structure of the workbook alone, and weak at anything requiring knowledge of what the numbers represent.

That's not a limitation of a particular vendor and it isn't going away with the next release. A spreadsheet carries its formulas and its layout, but it carries almost nothing about the world the formulas describe. A cell containing 0.0475 is a number; that it was meant to be this year's award increase and is now a year out of date is not information the file holds.

Structural
Determinable from the file itself — inconsistent formulas, broken ranges, hardcodes, circularity, orphaned cells. Reliably found, and found faster than by hand.
Semantic
Requires knowing what the number is for — a stale rate, a wrong funding assumption, a plausible figure in the wrong cell. Consistently missed.

What a model-audit pass reliably finds

Inconsistent formulas across a row or column. The single most valuable thing these tools do. Where 34 cells in a row share a formula and one doesn't, it gets surfaced immediately. This is the classic spreadsheet defect, it's genuinely hard to see by eye across a wide model, and it's where a real proportion of material errors live.

Hardcoded values buried inside formulas. The number typed into the middle of an expression during a late-night adjustment that then survives three years of reuse. These are found dependably, and the tool will list them, which by itself is a useful artefact.

Broken or truncated ranges. A sum covering rows 5 to 40 in a table that now runs to row 47. Easy for a machine, easy for a person to miss.

Structural questions about the model as a whole. Which sheets feed which, what the actual dependency chain is, where circular references sit, which inputs nothing depends on. Getting a written description of a workbook's architecture in a few minutes has genuine value when you've inherited a model and have no idea how it's wired.

What it reliably misses

A number that is plausible and wrong. A wage escalation set at 3.5 per cent when the applicable award movement is different, an occupancy assumption carried over from last year, a funding rate that changed in July. Every one of these is structurally perfect. Nothing in the file indicates a problem.

An assumption that's fine in isolation and wrong in combination. Revenue growth of 8 per cent alongside headcount held flat may be internally consistent arithmetic and operationally impossible. The tool checks the arithmetic.

The thing that isn't there. A model missing on-costs entirely, or omitting a leave provision movement, presents as clean. Absence has no formula.

Whether the model answers the question it was built for. The most consequential failure mode in practice, and completely outside the tool's reach.

The reason this division matters is that the errors in the second list are the expensive ones. A broken range usually produces a number so obviously wrong that someone notices. A stale escalation rate produces a number that looks right all the way to the board paper.

Why "audit my model" is the wrong prompt

Ask an assistant to audit a model and you get a mixed list: some real structural findings, some stylistic observations, some points that are simply wrong because the tool inferred intent it had no basis to infer. It reads authoritative, it takes twenty minutes to work through, and its main effect is to make you feel the model has been checked.

The problem isn't accuracy — it's that the request outsources the framing. You've asked an open question of something that will always produce an answer.

Narrower requests work much better because they play to the structural strength and don't invite the semantic guessing:

List every cell in this sheet whose formula differs from the cell to its left. List every hardcoded number appearing inside a formula, with its cell reference. Show me the full dependency chain feeding cell D48. List every input cell that nothing else in the workbook references.

Each returns a checkable list rather than an opinion. You can verify any line of it in seconds, which means you're reviewing evidence rather than trusting a conclusion — and that's the whole difference between a tool that helps and one that reassures.

Sensitivity analysis: the clearer win

Formula auditing is the advertised capability. Sensitivity work is where the time saving is larger and less discussed.

Asking for output across a range of values for two or three drivers, laid out as a table, is a task that has always been mechanical and has always been skipped under time pressure. Done quickly, it changes the conversation you can have — a board that sees a range with the breaking point marked is discussing risk, while a board that sees a single number is discussing whether they believe you.

One caution. The tool will happily flex a driver into territory that makes no operational sense, because it has no idea that your service can't scale past a certain point without another site. Bound the ranges yourself before asking.

Before a live workbook goes through an AI add-in: budget models routinely contain salaries by name, participant or client counts, and funder-specific rates. Confirm whether the vendor retains customer inputs for model training and prefer a tool where it doesn't — and check whether your organisation's tenancy settings match what the vendor's consumer product does, because they're often not the same arrangement. Better still, strip or de-identify names, employee identifiers and participant details before the workbook goes anywhere — the OAIC treats that minimisation as the primary control, not the vendor's terms.

A review workflow that holds up

Run the structural pass first, using narrow prompts, and fix what it finds. That clears the mechanical defects cheaply and puts you in front of a model that at least does what it appears to do.

Then spend what you just saved on the semantic review, by hand: every rate, every escalation, every date-dependent assumption, checked against its source, with the source noted next to it. In a typical model the list is short, because there usually aren't many. But it is the list carrying the risk, so it warrants more of your attention than the structural pass, not less — and it's the thing to hand a reviewer.

The tool takes the tedious half. The half that requires knowing what the organisation does stays where it always was.

When was the model behind your budget last independently reviewed?

Structural defects are now cheap to find. The assumptions still need someone who knows the sector. PFL provides senior-level outsourced finance, management reporting, and AI automation for Australian NFP, NDIS, and SME organisations.

Talk to PFL →
Timothy, CPA is Managing Director of Professional Financelink (PFL), providing senior-level outsourced finance, management reporting, and AI automation for Australian NFP, NDIS, and SME organisations. 20+ years in finance leadership across NFP, NDIS and SME.

Comments

Popular posts from this blog

Google Gemma 4 Just Launched — And It Might Solve Finance's Biggest AI Privacy Problem

Claude vs Gemini for Australian Finance: An Honest Comparison After 12 Months of Using Both

Why NFP Boards Are Finally Talking About AI — And What the Finance Team Should Do Before They Ask