Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

AutoSynthData turns agent failures into training tasks

ServiceNow reports improved enterprise task performance using generated training examples.

Listen to this article

ServiceNow’s CoreAI team described AutoSynthData on October 2, reporting better enterprise task performance after training a Gemma model on generated examples of work it previously handled poorly.

How the training loop works

The method uses a stronger teacher to turn weaknesses into new tasks, then checks successful solutions and deliberately incorrect outcomes before accepting samples. The generator receives capability summaries rather than original evaluation prompts.

Mean Pass@1 improved by 7.2 percentage points on EnterpriseOps Gym’s Hybrid tasks. On IT service management tasks it rose from 18.77 percent to 27.18 percent. The experiments used supervised fine tuning. Reinforcement learning remains future work.

Why the benchmark matters

The EnterpriseOps Gym paper was submitted in March. It describes a sandbox spanning eight enterprise domains, with 512 tools and 164 database tables. Hybrid tasks require work across domains. This existing benchmark is the setting for the new training results.

Its evaluation design checks the database state left after an agent acts. A successful workflow must achieve the requested result while respecting permissions and avoiding unwanted changes. Different action sequences can therefore pass. Completing an entire task is also distinct from satisfying only some of its checks.

The deployment question remains

For a business, that distinction matters. A support task might update the right ticket while also changing a record it should leave alone. A credible evaluation needs to catch both parts. Teams adopting an agent should define success in terms of the finished work and its side effects.

The reported gains support further testing in controlled environments. They do not establish reliability on a company’s live systems or an untouched evaluation set.

Before putting an agent in charge of a workflow, reserve separate tasks for evaluation and decide which mistakes still require human review. Test changes to permissions and records as carefully as the final answer. That makes a benchmark gain more useful to an operational decision.

Our Gemini 4 Argon evaluation coverage explains another reason to inspect the tasks behind an AI score before applying it elsewhere.

Featured image is an original AI generated illustration created for ByteForward. It depicts synthetic task validation conceptually.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile