All insights

Industry Copilot · 5 min read

Which Copilot Model Should Your Finance Team Actually Use?

By James Wilkinson 7 June 2026

Default for a quick email, but not for a forecast. When finance teams should switch the Copilot model to Claude Opus or GPT-5.5, keeping the P&L in-tenant.

TL;DR
  • For a quick finance email or a clean summary, Copilot's default model is the right tool. Leave the picker alone.
  • For a forecast with linked assumptions, structured variance work or anything going to the board or a lender, switch to GPT-5.5 or Claude Opus, then check the output.
  • Doing the analysis inside Copilot keeps the P&L in your governed Microsoft 365 tenant, instead of pasting figures into a personal ChatGPT account.

Finance teams were quick to find the useful corners of Copilot. Summarise a board pack, draft the commentary, pull last month’s figures into a note, tidy a reporting email. Most of that runs happily on whatever model Copilot uses by default, and you would not think to change it.

Then the work gets harder. A rolling forecast with linked assumptions. Variance commentary that has to trace back to the right driver. A reconciliation that will not behave. This is where the model picker, the one most teams never open, starts to matter.

The default is fine until the numbers get complicated

The default model is tuned for speed and routine work, and finance has plenty of that. For a quick email, a status update or a clean summary it is the right tool. Where it shows its limits is the multi-step reasoning a forecast demands: holding several driver cells in mind at once, following assumptions through to an output, explaining a movement rather than just restating it.

Ask the default to walk through a forecast with linked assumptions and you often get an answer that looks right and reads well. The trouble is the parts that look right and are not. In finance, that is the expensive kind of wrong. The same risk sits behind a lot of AI-assisted analysis and reporting in Excel: the output is only as trustworthy as the reasoning underneath it.

The failure mode is specific. The default rarely refuses or stalls. It produces a confident, fluent answer, and on a forecast that is the problem, because fluent and correct are not the same thing. A wrong number in a clean sentence is harder to catch than an obvious gap, and finance is the one place that error gets quoted onward as fact.

When to switch, and to which

The rule is the same one in the fuller guide to the Copilot model picker, applied to finance work.

  • Quick email, status update, clean summary: leave it on the default.
  • A forecast with linked drivers, a structured variance analysis, any reasoning that has to hold together: switch to GPT-5.5, built for analytical work, or to Claude Opus.
  • Anything that goes to the board, a lender or an auditor: switch, then check it line by line. Claude Opus is the one to reach for when you want the model to flag where it is unsure rather than present a tidy guess, which is the behaviour you want near a number you are about to sign off.

In practice the choice tracks the reporting cycle. The day-to-day chases, the holding emails and the quick figure-pulls sit on the default and stay there. The set pieces, month-end commentary, the reforecast, the pack that goes to the board, are the ones worth switching for. If your team only remembers one rule, make it that: anything someone outside finance will read closely, or make a decision on, deserves a switch and a second look.

Keep the numbers in-tenant

There is an old habit worth naming. Someone has a gnarly P&L question, the deadline is close, so they paste the figures into a personal ChatGPT account and get an answer in seconds. It works. It is also the firm’s financial data sitting in a consumer tool nobody approved.

The point of doing the same work inside Copilot is that the data does not leave. Switching to a more capable model in the picker keeps the analysis in your governed Microsoft 365 tenant, under the controls you already run. For accountancy practices handling client books the logic is sharper still, and it is the same discipline we set out for Copilot workflows in accountancy firms. A better answer is not worth much if you had to break a data rule to get it.

It is also a habit worth removing before it spreads. One person pasting a P&L into a consumer tool under deadline becomes the team’s unofficial method by the next quarter. Giving people a capable model inside Copilot is how you take that temptation away, rather than just warning against it.

A short worked example

Take a simple one. You ask Copilot to explain why Q3 gross margin fell against budget.

On the default, you tend to get a competent summary. Margin is down, here are the lines that moved, here is a sentence on each. Useful, quick, fine for an internal note.

Switch to a heavier model and give it the linked assumptions, and the answer changes shape. It traces the movement to the drivers: volume held, but input cost per unit rose and the price increase only landed part way through the quarter, so the margin effect is partial. It tells you which assumption to test next. That is the difference between a model that describes the variance and one that reasons about it.

Neither removes the review. You still check the figures yourself. But the second answer starts the conversation in a much better place.

It is worth being clear about what changed. The figures were the same in both runs. What improved was the reasoning over them, because the heavier model connected the assumptions instead of listing the lines. That is the gain you are buying when you switch, and it is largest on the questions where the answer is a judgement rather than a number.

Where to start

Most of this is judgement. Knowing which jobs deserve a switch and which do not. The dependable answer is not asking everyone to remember; it is wiring the right model into the workflow itself, so the heavier model sits behind the tasks that need it by default.

That is the work we do with finance teams and accountancy practices: agents and workflows built around the reports and reviews you already run. If you want the heavier models working on the right tasks, with the numbers kept in-tenant, book a free consultation and we will start from the reports your team runs each month.

Related reading

More on Industry Copilot

Industry Copilot Default, GPT-5.5 or Claude for Legal Document Review? A Law Firm's Guide to the Copilot Picker Which Copilot model for matter prep and contract review, plus the governance points UK law firms should settle before switching Claude on inside Copilot. Industry Copilot Five Admin Jobs an Agent Can Take Off a 30-Person Accountancy Firm Post, engagement letters, WIP reporting and meeting write-ups: five admin jobs we have built or scoped agents for, what stays with a person and what each costs. Industry Copilot How Accountancy Firms Actually Use Copilot In-House How UK accountancy firms use Microsoft Copilot in-house across client emails, reports and admin, and what makes the new habits last. Copilot Explainers The Copilot Model Picker Nobody Switches: When to Use Default, GPT-5.5 or Claude Opus Microsoft 365 Copilot has a model picker most teams never touch. When to use the default, GPT-5.5 or Claude Opus, and why you never leave Microsoft. Copilot Updates Claude Opus 5 in Microsoft 365 Copilot: Should UK Businesses Switch It On? Claude Opus 5 is live in the Microsoft 365 Copilot model selector. When it beats GPT-5.6, why UK tenants have it switched off and what enabling it involves. Copilot in Microsoft 365 Copilot in Excel: AI-Assisted Analysis, Reporting and Spreadsheet Automation Copilot in Excel can speed up analysis, formulas, variance commentary and reporting, but only when the workbook is clean and the question is specific. Industry Copilot Microsoft Copilot for Accountants: AI Workflows for Client Emails, Reports and Admin Microsoft Copilot for accountants: practical workflows for client emails, meeting summaries, report narratives and admin, with the governance UK firms need. Finance and reporting Hospitality P&L agent Monthly P&Ls arrived from each property in a different layout. The finance team now runs a Copilot Studio agent on demand to map each one onto the group's standard lines in seconds, and a person checks the output before the figures go anywhere. Industry AI agents for finance teams Agents that start when the month-end reports are exported, a budget meeting ends or a close deadline comes round, do the drafting and chasing, then stop for the finance team to review. Built in your own Microsoft 365. Industry AI agents for accountancy firms Six ready-made agents for the admin every practice shares: WIP and billing, client checks, meeting notes, records chasing, onboarding and VAT review. Each one waits for one of your team to approve. Next step Book a free consultation A free 30-minute call about the work an agent could take off your team, and whether Discover is the right next step.

Common questions

Questions about Copilot model for finance

Which Copilot model is best for forecasting?
For a forecast with linked assumptions, switch from the default to GPT-5.5 or Claude Opus. Both handle multi-step reasoning better than the default, which is tuned for routine work. Always check the output before you rely on it.
Is the default Copilot model good enough for finance work?
For quick emails, status updates and clean summaries, yes. For multi-driver forecasts, structured variance analysis or board-ready commentary, switch to a more capable model in the picker.
Can finance teams use Copilot without sending data to ChatGPT?
Yes. The models in the Copilot picker, including GPT-5.5 and Claude Opus, run inside your governed Microsoft 365 tenant. There is no need to paste a P&L into a personal ChatGPT or Claude.ai account.