Welcome to issue 3. AI can build a financial model with every formula right and still put the wrong number in it. This month's benchmark results put a figure on that, and it changes where your review time is best spent.

60-second skim
 
Sign off on numbers? The lead, then the CFO lens.
 
Build the models? What members built, then Try this.
 
Sign the deliverable? Straight to Regulation.
The main finding this month
Your review only tests the part that works
Vals AI asks a model to build a whole workbook from a brief and source files, roughly five hours of expert work. In results published 5 August, the best model passed:
88%
of the formula checks
80%
of the presentation checks
63%
of the numerical checks
Right formulas, right layout, wrong number. It is worst on LBO and DCF, where one early mistake flows through everything downstream. If you run a life-of-mine or life-of-field model, that is exactly this failure: a sound structure, a wrong price deck, and a defensible-looking waterfall on a wrong NPV.
"I think humans are actually not really great at checking either. Especially if it's going to be boring to check it, right? I think everybody just loses interest."
Gav, about forty minutes into the call
A formula passes as long as it points at the right cells. So if a cell holds the wrong number, the formula still works, still looks right, and still prints. The failures sit in the assumptions. Most review time goes on the structure.
"Even top scores are substantial-but-incomplete financial models, not client-ready deliverables."
Source · Vals AI, Excel Modeling Benchmark, updated 5 Aug 2026
01  ·  The CFO lens
Accuracy
Check the inputs, not the structure.
If your review traces the formulas and checks the formatting, it is testing the two things AI already gets right. Put that time on the inputs instead: the assumptions, and where each number came from.
Source · Vals AI, Excel Modeling Benchmark, 5 Aug 2026
Scope
Hand over steps you can check, not whole jobs.
Look a number up in a filing and calculate from it: 82%. Build the DCF or the LBO from these numbers: 35%. Same benchmark, same day, nine task types, best model in each. What predicts the score is how many steps sit between the question and the answer.
Source · Vals AI, Finance Agent v2, category leaders, updated 14 Aug 2026
Cost
Model price matters less than review time.
The runner-up scored within a rounding error of the winner for half the money, about US$6 a task against US$12. Check how your own AI subscription is billed too, because some plans will not let an outside tool draw on the allowance you have already paid for.
Source · Vals AI, Excel Modeling Benchmark, 5 Aug 2026

02  ·  What members built
Rough edges and all. Chatham House rules; named only with consent.
Build
A variance analyst you email a file to
Twenty to forty minutes later it emails back a CFO review of the variances and an executive deck. He had 80% of it working in a day. He calls it a management accountant analyst agent.
 
How he tested it. Fictional finance data seeded with deliberate traps, then a second model (Claude) grades the output and says how to improve it. That feedback goes back in. Five or six rounds so far.
 
How he keeps it honest. He tells it to stay general. Train it too hard on mining businesses, he says, and it goes blind on what other businesses have.
 
What went wrong. It acted without asking until he wrote it a persona saying you do this and nothing else. Charts came back with labels missing and series everywhere.
 
His verdict. It does not replace the analyst. It raises a red flag so the analyst knows where to dig.
Someone asked the obvious question: why not just paste the file into ChatGPT? For a one-off, he said, much the same result if you prompt it well. What the agent adds is that it has learned across many files how one person wants the analysis done, so nobody rebuilds the prompt every month.
Automation
The daily cash report, automated and monitored
Two people were spending two to three hours every morning logging into banks and platforms to build the daily cash and investment report. Now a robot does it at 5am on the company's own server, before anyone sits down. It has been running since February or March, and the hours it freed went into building the next robot. There are about ten now.
 
They chain. One robot downloads the invoices from the finance inbox. A second wakes when the first finishes, reads them, and posts the recurring ones straight into Xero.
 
Judgement is left alone. Utility bills yes, impairment no.
 
Everything is logged. He cannot read the logs himself and does not need to. When something breaks he throws the log at an AI and asks what happened. That folder is his audit trail: "If anyone, including if the auditors, ask where did all this number come from, these are the answer."
 
Failure is loud. A failed run never tries to fix itself, because an earlier attempt produced what he called "a really screwed up report". It goes red and waits for a person.
Once he passed three or four robots he lost track of which ones were working, so he built a dashboard: one card each, green or red, and the morning check takes ten seconds. During the demo, one card was red. Live, in front of everyone.

03  ·  Regulation
Regulation
The TPB says how the rules you already have cover AI.
The Tax Practitioners Board is a Commonwealth regulator, not a membership body. It regulates registered tax and BAS agents. Its July guidance sets out how the existing Code already applies:
 
You are accountable for the work, whoever produced it.
 
Client data going into an AI tool can need the client's permission first.
 
You owe reasonable care, and whether the work needs a second pair of eyes is part of what that requires.
 
You have to document your review.
Source · TPB(GS) 55/2026, 22 Jul 2026
Gav's read on the call: CPA Australia and CA ANZ will likely follow.

04  ·  What shipped since 7 July
The deck has the full list and every source.
 
Claude Opus 5 (24 Jul): close to Anthropic's best at half the price. Smarter and a lot more verbose. Cutting your own instructions back seems to help.
 
GPT-5.6 opened to everyone (9 Jul): three versions at three prices. OpenAI cut the cheapest by 80% on 30 July.
 
Gemini 3.6 Flash (21 Jul): competing on speed and price rather than raw ability. Handy for prototyping.
 
The Hugging Face incident (16 Jul): OpenAI's own models, running an internal cyber benchmark with the safety refusals turned down, got out. When Hugging Face investigated, the commercial services would not help, because their safety rules cannot tell an investigator from an attacker. It used an open-weights model already sitting on its own machines.
 
The EU AI Act took general effect (2 Aug): the heaviest paperwork is pushed back, with the high-risk obligations landing 2 December 2027. When they do, high-risk data points have to be documented and logged, which drags your vendors into the same discipline.

In the tools you already use: Xero data now works inside Claude, with Microsoft 365 Copilot support announced. NetSuite's assistant is US and Canada only. MYOB's AI BAS is beta and Australia only. With Power BI, Microsoft is reaching it through Copilot rather than building AI into the product.

05  ·  Try this

Give it a test it has to pass

In any AI tool, for anyone. When you ask it to look at a file, hand it the checks that have to pass and make it report them back with its answer:

 
The balance sheet balances.
 
The variance columns add to the movement.
 
The totals match the source file.

Each costs a line to write, and they catch the kind of mistake that reads perfectly well. A member asked whether you can build checks in. Gav's answer: put the validation instructions inside the instruction set, so the output arrives with the checks already ticked off. You do have to know what you want checked.

Try a workflow before you reach for an agent. An agent works out its own next step. A workflow follows the steps you wrote. A member made the case better than we can: for a month-end process, narrow and repeatable, you usually do not want an agent. You want a set prompt, deterministic checks, and the same result every time. Variability is the enemy. Gav's addition: chain the steps and the odds of failure compound.

If you are building something repeatable
Checks first· in the instructions, not in your review.
Set traps· fictional test files with deliberate mistakes, and a second model grading the output.
Narrow it· tell it what it does and nothing else.
Don't overfit· more than one industry, or it only finds the problems your last file had.
Fail loud· a failed run goes red and waits for a person.

Download the Meetup Presentation

Here’s a copy of the Meetup Presentation in case you weren’t able to make it. We’ve included recent news, as well as some tips and tricks.

260811 AI in Finance Meetup Final.pdf

260811 AI in Finance Meetup Final.pdf

1.54 MBPDF File

Join us next time

The next meetup

A monthly get-together for finance teams putting AI to work. News, a member demo, and breakout swaps on tips and the issues getting in your way. Chatham House rules, no spruiking, just finance people comparing notes.

date · 8th September 12:30pm AWST   |   Online · 60 min   |   Free

Save my seat →

Prepared in collaboration by  Model Answer × COD3R Lab

AI in Finance Meetup is a monthly community briefing for finance executives and teams working with AI. Recap drawn from the 11 August 2026 session. Demos and tips shared with members' consent; Chatham House rules apply. You're receiving this because you joined the community.

Model Answer, Perth WA, Australia. This is general information, not financial or investment advice.