Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Microsoft details limits in Foundry trace evaluation

Microsoft’s evaluation separates trace linkage from reliable agent diagnosis

Listen to this article

Microsoft published an October 2 evaluation report for Insights in Foundry, including a mixed AgentRx retail benchmark with mean trace precision of 37.8% and mean recall of 57.5%.

The tool entered public preview on September 24. It groups recurring agent behavior from production traces into findings with supporting evidence and suggested next steps.

What those figures measure

That benchmark contained 102 inputs, of which 29 were labeled failures. Precision measures the share of unique linked inputs bearing failure labels. Recall measures the share of labeled failures linked to a finding. Neither establishes that the diagnosis is correct.

Across six slices, Microsoft reused 746 inputs, including 653 labeled failures, on September 15, 16 and 17. Two slices contained only failures, making 100% precision automatic.

A separate September 23 controlled test detected ten of eleven scorable expected issues, with another issue unscored. A further evaluation used automated judges to assess findings. Those ratings were not human judgments and did not demonstrate that proposed fixes improved agents.

What a trial requires

The documentation requires a connected Application Insights resource, representative traces and a supported GPT 5 or newer judge deployment. Generation uses that deployment and may incur model charges, including scheduled scans.

Microsoft advises opening the supporting traces, comparing healthy examples and checking proposed changes before deployment. An empty result does not establish that an agent is healthy. The preview carries no service level agreement and is not recommended for production workloads.

For a team evaluating the service, a useful local test would mix known failures with successful requests from its own workflow. Record which findings a reviewer confirms, then test proposed changes against both groups. That connects a diagnostic suggestion to an actual release decision.

Archival Microsoft campus sign photographed in April 2005 by Derrick Coetzee, released into the public domain. This copy includes Zarex’s documented level adjustment. It illustrates Microsoft and does not show Foundry.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile