Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Microsoft tests cheaper agent evaluation with shared context

Microsoft tests composite agent evaluations while preview documentation details strict pass rules and limits on supported tools

Listen to this article

Microsoft published a study on October 2 testing whether one model call can judge several aspects of an AI agent. The approach shares conversation context across evaluation criteria instead of processing it repeatedly.

What the measurements show

On two separate 100 row samples, Microsoft reports input token reductions of 61.25% for response quality and 71.07% for tool use. Measured wall time fell 35.74% and 46.10%, respectively. Those figures describe the tested runs rather than guaranteed savings on a customer’s bill.

Quality varied by judge, rubric and workload. Groundedness was the least consistent dimension, so Microsoft recommends retaining focused fallback checks where needed.

How the preview handles failures

The Foundry documentation describes two composite evaluators. Output Quality combines six criteria, including task completion and groundedness. Tool Use Quality combines five checks covering tool selection, arguments, execution and use of returned results. Each keeps component scores and explanations.

The overall result passes only when every applicable component passes. A skipped component does not cause failure, and a run with every component skipped is marked not applicable. This makes the component details important when interpreting an apparently clean result.

Both evaluators are previews. Microsoft advises against production use for preview features. Its documentation also lists limited support for several tools, including Web Search and Code Interpreter, so teams should check their own tool mix before relying on the scores.

Keep the comparison honest

Start with agent conversations whose failures are already understood. Compare combined and separate judgments before changing an evaluation gate. A smaller token count is useful only if the checks still catch the mistakes that matter to the application.

Archival Microsoft campus photograph by Runner1928 from January 2015. Resized and converted to WebP. The image and this derivative are available under Creative Commons Attribution ShareAlike 4.0.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile