Copilot Agents · 7 min read
How Do You Know Your Agent Still Works?
By James Wilkinson 7 September 2026
Agents fail quietly. Copilot Studio's evaluation tools let you test an agent against a saved question set before and after every change. Most firms have never opened the tab.
TL;DR
- Agents do not fail with an error message. They keep answering, in the same confident tone, and the answers get worse. Nobody notices until a client does.
- Copilot Studio has had agent evaluation generally available since 31 March 2026. You build a test set of up to 100 questions, pick how answers are judged, save it and run it again after every change.
- The discipline is smaller than it sounds. Twenty questions written at build time, run before any change goes live, is most of the value.
- Three things to know first. The evaluation documentation covers agents on the standard harness, results are kept in Copilot Studio for 89 days, and AI-generated test cases can contain whatever data the test account can see.
- Evaluation measures correctness and performance. Microsoft is explicit that it does not measure AI safety or ethics, so it sits alongside your content filters and responsible AI review rather than replacing them.
Our Agent Governance Toolkit post made the case that an agent’s instructions are a request, not a control. This is the other half of that argument. Instructions are also not a guarantee. The agent that gave good answers in April can be giving poor ones in September, and nothing in your tenant will tell you.
The failure nobody notices
When a Power Automate flow breaks, it breaks. You get a failed run, an email and a red icon.
When an agent degrades, none of that happens. It still responds. It responds quickly, in the same tone, with the same confident structure. It has simply started leaving out the caveat, or citing last year’s policy, or answering a question the person did not ask. There is no error state for “worse”.
That is the whole problem in one paragraph. Every other system in your firm tells you when it stops working. This one does not.
Why agents drift
Nothing dramatic has to happen. The usual causes are mundane.
Someone edited the instructions. A partner asked for one change, a champion made it, and a sentence that was load-bearing got reworded.
The knowledge source moved on. The library the agent grounds on gained forty new documents, three of which are superseded drafts nobody deleted. The agent is now doing exactly what you told it, on worse source material.
The model changed. Microsoft updates and retires the models behind agents on its own schedule. An instruction that was interpreted one way in the spring can be interpreted differently later.
Tools were added. A new connector went in for a different use case. The agent now has an extra option and sometimes reaches for it at the wrong moment.
The questions changed. The agent was built for the questions people asked in month one. By month six they are asking different things, and the design has quietly drifted away from the demand.
What Copilot Studio actually gives you
There is an Evaluation tab on every agent, and in our experience almost nobody has opened it. Agent evaluation went generally available on 31 March 2026.
You create a test set, which is a group of up to 100 test cases. You can write them yourself, import them from a file, generate them or pull them from real conversations. Copilot Studio runs every case in the set against your agent and reports the results.
The generation options are the quickest way in. A quick question set produces ten questions automatically from the agent’s description, instructions and capabilities, which is enough for a fast sanity check. A full question set generates questions from a knowledge source or from the agent’s topics, and you choose how many. Generation supports text, Word, Excel, PDF and individual SharePoint files up to 5 MB. SharePoint folders are not supported.
Importing is deliberately plain. The file is a CSV or text file with two columns, Question and Expected response, up to 100 questions, each question up to 1,000 characters. Expected responses are optional on import, but the match, similarity and compare meaning methods need them.
You can also build a test set from real user questions. Copilot Studio groups questions from actual conversations into themes, currently in preview, and you can turn a theme straight into an evaluation. That is the most honest test set you will ever have, because you did not write it.
Alongside single response tests there are conversational evaluations, which assess how the agent holds up across a longer exchange rather than one question at a time. Most real failures happen on turn three, not turn one.
How the answers get judged
Copilot Studio calls these test methods. There are seven, and you can combine several in one test set.
General quality is the default and needs no configuration. It uses a model to score the response against four criteria: relevance, groundedness, completeness and abstention. Groundedness is the one that matters most for professional services, because it asks whether the answer is actually supported by the retrieved sources rather than invented around them. A response has to meet all four criteria to be considered high quality.
Compare meaning checks whether the answer carries the intended meaning of an expected response, without requiring the same wording. The default pass score is 50.
Tool use checks whether the agent used the tools or topics you expected. This is how you catch an agent that gets to a plausible answer by the wrong route.
Keyword match, text similarity and exact match handle the cases where specific wording matters. Keyword match lets you require any or all of your keywords. Text similarity is the one to reach for when precise phrasing is required, such as generated legal wording, and Microsoft’s own documentation gives that as the example.
Custom is the one that matters most for a firm, and the one almost nobody configures. You write evaluation instructions in plain English describing what you want checked, then define labels with pass or fail assignments. Microsoft’s own example is an HR compliance test labelling answers compliant or non-compliant.
That is where your actual standards go. Not “is this a good answer” but “does this cite the current engagement terms”, “does this decline to give an opinion on the client’s tax position” or “does this include the disclaimer”.
The part that makes it a control rather than an exercise
Test sets are saved and reused. Run the same set twice and the Compare with tool puts the two runs against each other, with arrows on every case that moved from failing to passing or from passing to failing.
That is the difference between testing an agent once and having a control. One run tells you the agent works today. The same run after a change tells you whether the change helped or quietly broke something. Copilot Studio also records response time for each interaction, which is a measurement only and does not affect the pass rate, but it is how you spot an agent that got slower rather than wronger.
You can also trigger evaluations through the Power Platform REST APIs or through connectors added as tools or as steps in a Copilot Studio or Power Automate flow, so the check runs without anyone remembering to run it. Only one evaluation runs at a time, so a hundred-case set is a few minutes rather than a background job you forget about.
Four things to know before you rely on it
Check what your agent is built on. Microsoft’s evaluation documentation is marked as describing features on the standard harness. Copilot Studio has three harnesses now, the GitHub Copilot harness, the standard harness and the Copilot chat harness, and the GitHub Copilot harness became the default for new builds on 3 August 2026. An agent cannot move between harnesses later, so this is worth checking before you plan a testing process around it. Our post on choosing a harness covers the difference and what the change costs.
Results are kept for 89 days. After that they are gone from Copilot Studio. If you need evidence of testing for an audit, a client or your own quality records, export the results to CSV and keep them somewhere permanent. For regulated firms this is not an optional step.
Generated test cases can contain sensitive data. Evaluations run under a selected test account and use that account’s authentication to reach knowledge sources and tools. Anything that account can see can end up in a generated test case, and any maker with access to the agent can view the test sets. Choose the test account deliberately, the same way you would choose a service account. The same thinking applies as in our guide to using Copilot with client data.
It is not a safety check. Microsoft says plainly that agent evaluation measures correctness and performance, not AI ethics or safety. An agent can pass every case in your set and still give an answer you would not want sent to a client. Content filters and human review points stay where they are. Our post on where agent data goes covers the other half of that.
One smaller limitation worth knowing: agent evaluation does not currently support Fabric data agents.
What a firm of forty people should actually do
You do not need a QA function. You need a habit.
- Write twenty questions at build time. The moment to do this is while you still remember what a good answer looks like. Include the awkward ones: the question the agent should refuse, the one where the honest answer is “ask a person”, the one where an out of date document would produce a plausible wrong answer.
- Add one custom test. One rule that matters to your firm, written in plain English. That single test is worth more than the other six methods combined, because it encodes your standard rather than a generic one.
- Run it before any change goes live. Instruction edits, knowledge source changes, new tools. If the pass rate drops, you know before your clients do.
- Run it quarterly regardless. Things change underneath you that nobody in the firm touched.
- Name who owns it. An unowned check is not a check. There is a practical wrinkle here: any maker can see the pass rate of a run, but only the maker who started a run can see the agent’s actual answers and the reasoning behind each result. If the person who runs the check is not the person who reads the failures, decide now which of them it is. Copilot Studio also has an agent viewer sharing role, which gives someone the Evaluation tab without giving them the agent.
That is the fourth question to add to the three in our governance post: what exactly can it do, which agent did what, where is the record, and who runs the check.
Roughly half a day at build time and an hour a quarter after that.
Agents are not furniture
The uncomfortable truth about agents is that they are not something you install and forget. They sit on top of your content, your permissions, your instructions and a model that Microsoft updates without asking you. All four move. It is the same argument we make about automation that nobody maintains, with a system that hides its failures better. What happens after the agent is built sets out the five-point minimum an agent’s owner runs, and the saved question set is point five of it.
That is exactly what our Embed stage is for: the agents keep working because someone is checking that they do. The Agent Journey shows where that sits alongside designing and building agents, and the agents page covers what we build. If you have agents live in your tenant and no one has opened the Evaluation tab, book a free 30 minute consultation and we will look at what you have running and what a sensible test set for it would contain.
Sources checked
Last checked: 7 September 2026.
- Microsoft Learn, “About agent evaluation”
- Microsoft Learn, “Choose evaluation methods”
- Microsoft Learn, “Create a single response test set”
- Microsoft Learn, “Run evaluations and view results”
- Microsoft Learn, “Harnesses in Copilot Studio”
- Microsoft Learn, “Choose between Agent Builder in Microsoft 365 Copilot and Copilot Studio”
- Microsoft Copilot Studio Blog, “Agent Evaluation in Microsoft Copilot Studio is now generally available”, 31 March 2026
Related reading
More on Copilot Agents
Common questions