Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Why METR time horizons need more context

Listen to this article

A Berkeley preprint submitted October 8 asks whether equal multipliers in METR time horizons represent equal capability gains. Drew T. Nguyen and William Fithian reanalyzed 228 tasks and 26 AI systems.

What the number measures

METR measures task length by the time a human expert needs to finish it. An AI system’s 50 percent horizon is the duration at which its fitted success curve reaches an even chance of completion. The number describes performance on the benchmark’s task distribution. Individual tasks of similar duration can produce very different outcomes. It is also separate from the time the AI itself spends running.

The benchmark mainly covers software engineering, machine learning and cybersecurity. Its tasks have clear instructions and automated scoring. METR says the human comparison is closer to a skilled newcomer with little project context than an employee who already knows the codebase. Those details belong beside any headline horizon figure. METR marked the chart as no longer actively updated on September 8.

Why equal jumps can differ

The baseline assumes a particular smooth decline in success as human task duration grows. More precisely, success log odds are linear in log duration, with a common slope across systems. Log odds compare success with failure on a logarithmic scale.

Nguyen and Fithian fit a shared flexible curve, with one version allowing for task family effects. Both improve point estimates across 12 validation metrics using withheld task families.

The fitted difficulty curve is nearly flat from 2 to 30 minutes. Equal multipliers can therefore represent different capability gains. The authors retain the observed exponential trend and recommend diagnostic plots.

Consider a hypothetical success curve that falls linearly from 52 percent at 5 minutes to 48 percent at 25 minutes. Raising it by two percentage points moves the 50 percent horizon from 15 to 25 minutes.

The diagnostics assume representative tasks, meaningful human timing and adequate single dimensional summaries of ability and difficulty.

The earlier warnings from METR

In a January 22 research note, METR researcher Thomas Kwa had already warned against reading a 50 percent horizon as permission to delegate every shorter task. Some work needs much higher reliability. Converting a horizon into a productivity gain also requires accounting for prompting, checking and recovery effort. Kwa stressed that uncertainty and differences between task domains limit what a precise looking number can tell readers.

Alexander Barry’s March 20 sensitivity analysis explored alternative fits and uncertainty in human timing. He found that reasonable choices generally stayed inside the already wide confidence intervals, while the task distribution remained a major source of uncertainty. A logistic model using one common slope performed best on his two validation measures. These historical notes predate the Berkeley paper and are individual research updates rather than a new institutional response.

Questions to ask before deployment

For a team choosing an AI system, the practical next step is to specify an acceptable outcome before comparing headline scores. A draft that an engineer can check quickly and an unattended change to a production system call for different evidence. Record the cost of a failed attempt, who will detect it and how much work recovery requires.

Then test a representative sample of the intended workload. Keep the instructions, tools and completion criteria consistent across candidates. Record successes alongside human review time and unresolved errors.

Evans Hall from Sather Tower. UC Berkeley campus in May 2022. Campus context for the Berkeley research. Photograph by Gabe Classon, CC BY 2.0. Original photograph.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile