All insights

Copilot Agents · 7 min read

How Do You Know Your Agent Still Works?

By James Wilkinson 7 September 2026

Agents fail quietly. Copilot Studio's evaluation tools let you test an agent against a saved question set before and after every change. Most firms have never opened the tab.

TL;DR
  • Agents do not fail with an error message. They keep answering, in the same confident tone, and the answers get worse. Nobody notices until a client does.
  • Copilot Studio has had agent evaluation generally available since 31 March 2026. You build a test set of up to 100 questions, pick how answers are judged, save it and run it again after every change.
  • The discipline is smaller than it sounds. Twenty questions written at build time, run before any change goes live, is most of the value.
  • Three things to know first. The evaluation documentation covers agents on the standard harness, results are kept in Copilot Studio for 89 days, and AI-generated test cases can contain whatever data the test account can see.
  • Evaluation measures correctness and performance. Microsoft is explicit that it does not measure AI safety or ethics, so it sits alongside your content filters and responsible AI review rather than replacing them.

Our Agent Governance Toolkit post made the case that an agent’s instructions are a request, not a control. This is the other half of that argument. Instructions are also not a guarantee. The agent that gave good answers in April can be giving poor ones in September, and nothing in your tenant will tell you.

The failure nobody notices

When a Power Automate flow breaks, it breaks. You get a failed run, an email and a red icon.

When an agent degrades, none of that happens. It still responds. It responds quickly, in the same tone, with the same confident structure. It has simply started leaving out the caveat, or citing last year’s policy, or answering a question the person did not ask. There is no error state for “worse”.

That is the whole problem in one paragraph. Every other system in your firm tells you when it stops working. This one does not.

Why agents drift

Nothing dramatic has to happen. The usual causes are mundane.

Someone edited the instructions. A partner asked for one change, a champion made it, and a sentence that was load-bearing got reworded.

The knowledge source moved on. The library the agent grounds on gained forty new documents, three of which are superseded drafts nobody deleted. The agent is now doing exactly what you told it, on worse source material.

The model changed. Microsoft updates and retires the models behind agents on its own schedule. An instruction that was interpreted one way in the spring can be interpreted differently later.

Tools were added. A new connector went in for a different use case. The agent now has an extra option and sometimes reaches for it at the wrong moment.

The questions changed. The agent was built for the questions people asked in month one. By month six they are asking different things, and the design has quietly drifted away from the demand.

What Copilot Studio actually gives you

There is an Evaluation tab on every agent, and in our experience almost nobody has opened it. Agent evaluation went generally available on 31 March 2026.

You create a test set, which is a group of up to 100 test cases. You can write them yourself, import them from a file, generate them or pull them from real conversations. Copilot Studio runs every case in the set against your agent and reports the results.

The generation options are the quickest way in. A quick question set produces ten questions automatically from the agent’s description, instructions and capabilities, which is enough for a fast sanity check. A full question set generates questions from a knowledge source or from the agent’s topics, and you choose how many. Generation supports text, Word, Excel, PDF and individual SharePoint files up to 5 MB. SharePoint folders are not supported.

Importing is deliberately plain. The file is a CSV or text file with two columns, Question and Expected response, up to 100 questions, each question up to 1,000 characters. Expected responses are optional on import, but the match, similarity and compare meaning methods need them.

You can also build a test set from real user questions. Copilot Studio groups questions from actual conversations into themes, currently in preview, and you can turn a theme straight into an evaluation. That is the most honest test set you will ever have, because you did not write it.

Alongside single response tests there are conversational evaluations, which assess how the agent holds up across a longer exchange rather than one question at a time. Most real failures happen on turn three, not turn one.

How the answers get judged

Copilot Studio calls these test methods. There are seven, and you can combine several in one test set.

General quality is the default and needs no configuration. It uses a model to score the response against four criteria: relevance, groundedness, completeness and abstention. Groundedness is the one that matters most for professional services, because it asks whether the answer is actually supported by the retrieved sources rather than invented around them. A response has to meet all four criteria to be considered high quality.

Compare meaning checks whether the answer carries the intended meaning of an expected response, without requiring the same wording. The default pass score is 50.

Tool use checks whether the agent used the tools or topics you expected. This is how you catch an agent that gets to a plausible answer by the wrong route.

Keyword match, text similarity and exact match handle the cases where specific wording matters. Keyword match lets you require any or all of your keywords. Text similarity is the one to reach for when precise phrasing is required, such as generated legal wording, and Microsoft’s own documentation gives that as the example.

Custom is the one that matters most for a firm, and the one almost nobody configures. You write evaluation instructions in plain English describing what you want checked, then define labels with pass or fail assignments. Microsoft’s own example is an HR compliance test labelling answers compliant or non-compliant.

That is where your actual standards go. Not “is this a good answer” but “does this cite the current engagement terms”, “does this decline to give an opinion on the client’s tax position” or “does this include the disclaimer”.

The part that makes it a control rather than an exercise

Test sets are saved and reused. Run the same set twice and the Compare with tool puts the two runs against each other, with arrows on every case that moved from failing to passing or from passing to failing.

That is the difference between testing an agent once and having a control. One run tells you the agent works today. The same run after a change tells you whether the change helped or quietly broke something. Copilot Studio also records response time for each interaction, which is a measurement only and does not affect the pass rate, but it is how you spot an agent that got slower rather than wronger.

You can also trigger evaluations through the Power Platform REST APIs or through connectors added as tools or as steps in a Copilot Studio or Power Automate flow, so the check runs without anyone remembering to run it. Only one evaluation runs at a time, so a hundred-case set is a few minutes rather than a background job you forget about.

Four things to know before you rely on it

Check what your agent is built on. Microsoft’s evaluation documentation is marked as describing features on the standard harness. Copilot Studio has three harnesses now, the GitHub Copilot harness, the standard harness and the Copilot chat harness, and the GitHub Copilot harness became the default for new builds on 3 August 2026. An agent cannot move between harnesses later, so this is worth checking before you plan a testing process around it. Our post on choosing a harness covers the difference and what the change costs.

Results are kept for 89 days. After that they are gone from Copilot Studio. If you need evidence of testing for an audit, a client or your own quality records, export the results to CSV and keep them somewhere permanent. For regulated firms this is not an optional step.

Generated test cases can contain sensitive data. Evaluations run under a selected test account and use that account’s authentication to reach knowledge sources and tools. Anything that account can see can end up in a generated test case, and any maker with access to the agent can view the test sets. Choose the test account deliberately, the same way you would choose a service account. The same thinking applies as in our guide to using Copilot with client data.

It is not a safety check. Microsoft says plainly that agent evaluation measures correctness and performance, not AI ethics or safety. An agent can pass every case in your set and still give an answer you would not want sent to a client. Content filters and human review points stay where they are. Our post on where agent data goes covers the other half of that.

One smaller limitation worth knowing: agent evaluation does not currently support Fabric data agents.

What a firm of forty people should actually do

You do not need a QA function. You need a habit.

  1. Write twenty questions at build time. The moment to do this is while you still remember what a good answer looks like. Include the awkward ones: the question the agent should refuse, the one where the honest answer is “ask a person”, the one where an out of date document would produce a plausible wrong answer.
  2. Add one custom test. One rule that matters to your firm, written in plain English. That single test is worth more than the other six methods combined, because it encodes your standard rather than a generic one.
  3. Run it before any change goes live. Instruction edits, knowledge source changes, new tools. If the pass rate drops, you know before your clients do.
  4. Run it quarterly regardless. Things change underneath you that nobody in the firm touched.
  5. Name who owns it. An unowned check is not a check. There is a practical wrinkle here: any maker can see the pass rate of a run, but only the maker who started a run can see the agent’s actual answers and the reasoning behind each result. If the person who runs the check is not the person who reads the failures, decide now which of them it is. Copilot Studio also has an agent viewer sharing role, which gives someone the Evaluation tab without giving them the agent.

That is the fourth question to add to the three in our governance post: what exactly can it do, which agent did what, where is the record, and who runs the check.

Roughly half a day at build time and an hour a quarter after that.

Agents are not furniture

The uncomfortable truth about agents is that they are not something you install and forget. They sit on top of your content, your permissions, your instructions and a model that Microsoft updates without asking you. All four move. It is the same argument we make about automation that nobody maintains, with a system that hides its failures better. What happens after the agent is built sets out the five-point minimum an agent’s owner runs, and the saved question set is point five of it.

That is exactly what our Embed stage is for: the agents keep working because someone is checking that they do. The Agent Journey shows where that sits alongside designing and building agents, and the agents page covers what we build. If you have agents live in your tenant and no one has opened the Evaluation tab, book a free 30 minute consultation and we will look at what you have running and what a sensible test set for it would contain.

Sources checked

Last checked: 7 September 2026.

Related reading

More on Copilot Agents

Copilot Agents What Happens After the Agent Is Built? Why Every Agent Needs an Owner An agent is a running service, not a finished project. What changes after go-live, the five-point minimum a firm with no IT team can run and who owns it. Copilot Agents Microsoft's Agent Governance Toolkit: Three Questions to Ask of Any Agent Microsoft's open-source Agent Governance Toolkit enforces agent rules in code, not prompts. Why that matters for firms that will never install it. Copilot Agents What Is a Copilot Agent? A Plain-English Definition A plain-English definition of Microsoft Copilot agents: what they are made of, the jobs they do well, where they run and how firms go from one agent to a team. Automation Which Copilot Studio Harness Should Your Agent Be Built On? An agent harness explained in plain English, why the GitHub Copilot harness costs more and how to decide which harness each agent build belongs on. Automation Copilot Studio's New Engine: What the GitHub Copilot Harness Changes and What It Costs Copilot Studio's default is now the GitHub Copilot harness, the engine behind Copilot Cowork. What changes, what the credits cover and the checks before you build. Copilot Agents Are Copilot Agents Secure? Where Your Data Goes The long answer on Copilot agent security: where answers come from, what stays inside the Microsoft 365 service boundary and what your admins control. Copilot Governance Can You Use Microsoft Copilot with Client Data? A Practical Governance Guide Can you use Microsoft Copilot with client data? Yes, but only inside clear governance, approved tools and a review process built around risk levels. Automation Automation Maintenance for Microsoft 365: Why Workflows Need Ongoing Support Automation maintenance is what keeps Microsoft 365 workflows saving time. Owners, monitoring, documentation and change reviews stop quiet failures. Service Copilot Studio consultancy Agents that start on their own, act in your systems and stop for one of your team to approve, built in Copilot Studio inside your Microsoft 365. Service area Embed We keep your agents running. Capability map The Agent Journey Five stages from using your first Copilot agent to running a team of agents, with two doors at every stage. Agents Microsoft Copilot agents Plain-English guidance on where Microsoft Copilot agents fit, how to govern them and when to build. Next step Book a free consultation A free 30-minute call about the work an agent could take off your team, and whether Discover is the right next step.

Common questions

Questions about testing Copilot Studio agents

Do we need extra licences to evaluate a Copilot Studio agent?
Evaluation is built into Copilot Studio rather than sold separately, so there is no additional tool to buy. Each test case is a real request to the agent, so a hundred-case test set is a hundred agent responses, billed the same way that agent's ordinary traffic is billed. On the standard harness that means the Copilot Credit rate card once the agent is published. On the GitHub Copilot harness the meter runs while you build and test as well. Microsoft does not publish a separate rate for evaluation.
How many test questions do we need?
A test set holds up to 100 cases, but twenty well-chosen questions beat a hundred generated ones. Include the questions the agent should refuse, not just the ones it should answer well.
Can we test an agent automatically instead of remembering to run it?
Yes. Evaluations can be triggered through the Power Platform REST APIs or through connectors added as tools or as steps in a Copilot Studio or Power Automate flow, so the check can run on a schedule or as part of a release process. Only one evaluation runs at a time on a given agent.
How long are the results kept?
Test results stay in Copilot Studio for 89 days. Export them to CSV if you need a longer record, which most regulated firms will. The export includes the question, expected response, test method, pass score, the agent's answer, the result and the analysis for each case.
Does agent evaluation tell us whether an agent is safe?
No. Microsoft states that agent evaluation measures correctness and performance rather than AI ethics or safety problems. An agent can pass every test in a set and still produce an inappropriate answer, so content filters and responsible AI review stay in place alongside it.
What if our agent was built in Microsoft 365 Copilot rather than Copilot Studio?
The evaluation tooling described here lives in Copilot Studio. Agents created with Agent Builder in Microsoft 365 Copilot can be copied into Copilot Studio, which is also how you reach multistep workflows, custom integrations and broader deployment options.