Industry Copilot · 5 min read
Which Copilot Model Should Your Finance Team Actually Use?
By James Wilkinson 7 June 2026
Default for a quick email, but not for a forecast. When finance teams should switch the Copilot model to Claude Opus or GPT-5.5, keeping the P&L in-tenant.
TL;DR
- For a quick finance email or a clean summary, Copilot's default model is the right tool. Leave the picker alone.
- For a forecast with linked assumptions, structured variance work or anything going to the board or a lender, switch to GPT-5.5 or Claude Opus, then check the output.
- Doing the analysis inside Copilot keeps the P&L in your governed Microsoft 365 tenant, instead of pasting figures into a personal ChatGPT account.
Finance teams were quick to find the useful corners of Copilot. Summarise a board pack, draft the commentary, pull last month’s figures into a note, tidy a reporting email. Most of that runs happily on whatever model Copilot uses by default, and you would not think to change it.
Then the work gets harder. A rolling forecast with linked assumptions. Variance commentary that has to trace back to the right driver. A reconciliation that will not behave. This is where the model picker, the one most teams never open, starts to matter.
The default is fine until the numbers get complicated
The default model is tuned for speed and routine work, and finance has plenty of that. For a quick email, a status update or a clean summary it is the right tool. Where it shows its limits is the multi-step reasoning a forecast demands: holding several driver cells in mind at once, following assumptions through to an output, explaining a movement rather than just restating it.
Ask the default to walk through a forecast with linked assumptions and you often get an answer that looks right and reads well. The trouble is the parts that look right and are not. In finance, that is the expensive kind of wrong. The same risk sits behind a lot of AI-assisted analysis and reporting in Excel: the output is only as trustworthy as the reasoning underneath it.
The failure mode is specific. The default rarely refuses or stalls. It produces a confident, fluent answer, and on a forecast that is the problem, because fluent and correct are not the same thing. A wrong number in a clean sentence is harder to catch than an obvious gap, and finance is the one place that error gets quoted onward as fact.
When to switch, and to which
The rule is the same one in the fuller guide to the Copilot model picker, applied to finance work.
- Quick email, status update, clean summary: leave it on the default.
- A forecast with linked drivers, a structured variance analysis, any reasoning that has to hold together: switch to GPT-5.5, built for analytical work, or to Claude Opus.
- Anything that goes to the board, a lender or an auditor: switch, then check it line by line. Claude Opus is the one to reach for when you want the model to flag where it is unsure rather than present a tidy guess, which is the behaviour you want near a number you are about to sign off.
In practice the choice tracks the reporting cycle. The day-to-day chases, the holding emails and the quick figure-pulls sit on the default and stay there. The set pieces, month-end commentary, the reforecast, the pack that goes to the board, are the ones worth switching for. If your team only remembers one rule, make it that: anything someone outside finance will read closely, or make a decision on, deserves a switch and a second look.
Keep the numbers in-tenant
There is an old habit worth naming. Someone has a gnarly P&L question, the deadline is close, so they paste the figures into a personal ChatGPT account and get an answer in seconds. It works. It is also the firm’s financial data sitting in a consumer tool nobody approved.
The point of doing the same work inside Copilot is that the data does not leave. Switching to a more capable model in the picker keeps the analysis in your governed Microsoft 365 tenant, under the controls you already run. For accountancy practices handling client books the logic is sharper still, and it is the same discipline we set out for Copilot workflows in accountancy firms. A better answer is not worth much if you had to break a data rule to get it.
It is also a habit worth removing before it spreads. One person pasting a P&L into a consumer tool under deadline becomes the team’s unofficial method by the next quarter. Giving people a capable model inside Copilot is how you take that temptation away, rather than just warning against it.
A short worked example
Take a simple one. You ask Copilot to explain why Q3 gross margin fell against budget.
On the default, you tend to get a competent summary. Margin is down, here are the lines that moved, here is a sentence on each. Useful, quick, fine for an internal note.
Switch to a heavier model and give it the linked assumptions, and the answer changes shape. It traces the movement to the drivers: volume held, but input cost per unit rose and the price increase only landed part way through the quarter, so the margin effect is partial. It tells you which assumption to test next. That is the difference between a model that describes the variance and one that reasons about it.
Neither removes the review. You still check the figures yourself. But the second answer starts the conversation in a much better place.
It is worth being clear about what changed. The figures were the same in both runs. What improved was the reasoning over them, because the heavier model connected the assumptions instead of listing the lines. That is the gain you are buying when you switch, and it is largest on the questions where the answer is a judgement rather than a number.
Where to start
Most of this is judgement. Knowing which jobs deserve a switch and which do not. The dependable answer is not asking everyone to remember; it is wiring the right model into the workflow itself, so the heavier model sits behind the tasks that need it by default.
That is the work we do with finance teams and accountancy practices: agents and workflows built around the reports and reviews you already run. If you want the heavier models working on the right tasks, with the numbers kept in-tenant, book a free consultation and we will start from the reports your team runs each month.
Related reading
More on Industry Copilot
Common questions