Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Replitโ€™s coding agents solved more tasks with flexible delegation

More flexible delegation raised benchmark scores and the cost per task.

Listen to this article

Replitโ€™s coding agent solved more benchmark tasks when it could choose how to delegate. Costs also rose. In the September 29 report, GPT 6 Astra chooses workers, their size and reasoning effort as a task unfolds.

The gains came with a higher bill

Against one persistent worker, Replit reports scores rising from 61 to 72 percent on DeepSWE 1.1 and 33 to 49 percent on Terminal Bench 4.0. Reported cost per task also rose, from $1.34 to $2.11 and $1.84 to $2.53 respectively.

The comparison covered a selected set of tasks

The company averaged four repetitions across 113 DeepSWE tasks and 63 Terminal Bench tasks. The Terminal Bench run excluded three GPU tasks. The comparison changes the delegation architecture while retaining the configuration. The results have not been independently reproduced for this report.

The Terminal Bench organizers explain that task revisions and resource settings can change results enough to require new trials. Their full benchmark includes GPU tasks. That makes the exact task selection, execution environment and agent setup important when comparing Replitโ€™s reported subset with results elsewhere.

Model routing was already in use

Replit introduced automatic model routing on August 26. That announcement described matching models to evolving tasks, with user overrides and enterprise controls over the available model set. The September report adds a closer look at delegation and its measured tradeoffs. The report includes a September 17 production trace, placing this behavior before the research publication.

How much is another completed task worth

For an engineering team, the decision comes down to the work those extra dollars buy. Compare both approaches on the same tasks within a set budget, then count completed work, failed attempts and the time spent reviewing results. A benchmark gain is useful only if it carries over to the teamโ€™s own workload.

Illustrative archival photograph of a coding workstation by Farzad Nazifi, dated 2016, via Wikimedia Commons under CC0 1.0. It does not show Replit or its benchmark experiments. No affiliation or endorsement is implied.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile