Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Independent Gemini 4 Argon tests show strong scores and uneven agent skills

Independent Gemini 4 Argon tests show a strong Vals Index result and a close Astra comparison, with important pricing and game control limits.

Listen to this article

Independent evaluations published on September 30 give Gemini 4 Argon strong results across professional tasks, alongside a much weaker showing in a video game control test. The findings from Vals AI and Artificial Analysis add evidence beyond Google’s own launch comparisons.

Argon takes the lead on the Vals Index

Vals AI’s model report lists Argon first among 41 models on its Vals Index, with a score of 68.90 percent and a reported standard error of 0.97 percentage points. Its listed evaluation settings include high reasoning effort and a temperature of 1, with individual benchmarks able to use different parameters.

Four model Vals Index scores and test costs with Gemini 4 Argon at 68.90 percent and 15.68 dollars per test
Chart by ByteForward using the Vals AI September 30 report. Cost is specific to the test. Small score gaps do not establish statistical significance.
ModelVals Index scoreCost per test in USD
Gemini 4 Argon68.90%15.68
Claude Sonnet 5.567.04%21.34
Claude Opus 5.566.97%32.14
Claude Fable 5.165.83%28.71
Data behind the chart from the original evaluator

The index methodology combines eight benchmarks covering finance, coding, legal and tax work. Sector weights reflect shares of the US economy, so finance and coding contribute more than the smaller categories. This is a weighted test score, not a measured increase in economic output or a percentage of jobs that can be automated.

Vals reports standard errors to describe statistical uncertainty. Those error bars do not capture every change a different prompt, software setup or deployment could produce. A leading position is evidence about the evaluated tasks and settings.

Artificial Analysis finds a close match with Astra

In its September 30 evaluation, Artificial Analysis reports a rounded Intelligence Index score of 53 for Argon at high reasoning, matching GPT 6 Astra at maximum reasoning. That is a different measure from the Vals Index, so the two headline numbers cannot be compared directly.

Artificial Analysis’ current methodology combines ten evaluations across agents, coding, scientific reasoning and general capabilities. It is primarily an English language, text based suite. Visual, speech and multilingual performance are evaluated separately, which limits what the overall score says about those capabilities.

The firm estimates Argon’s cost at $1.99 per Intelligence Index task under promotional pricing, versus $3.26 for Astra. That is about 60 percent of Astra’s cost, or roughly 40 percent less. At the announced standard prices, it estimates Argon would cost $3.98 per task.

Artificial Analysis attributes the advantage to lower token prices rather than fewer generated tokens. Its Argon runs averaged about 62,000 output tokens per task, compared with 27,000 for Astra. The discount therefore matters materially to the comparison.

Vals lists $15.68 per Vals Index test using input and output rates of $4 and $20 per million tokens. Artificial Analysis uses the introductory rates of $2 and $10. Different task collections and pricing assumptions mean these dollar figures are not interchangeable estimates of the same workload.

The game control result measures a different problem

Vals reports Argon at 4.83 out of 100 on its separate CUA benchmark, seventh among eight models. The test methodology involves six commercial video games, with three hours allowed per game. Agents see screenshots and act through keyboard and mouse controls. Half the games are undisclosed. Agents can also use scratch tools, and the published comparison has one trial per model and game.

Scores average progress through game milestones. Response delays are part of the challenge because the games can keep moving while an agent decides what to do. The result identifies a weakness in this particular control setting. It does not establish a failure rate for ordinary office software or every computer task.

What the independent results establish

Together, the studies support treating Argon as a serious candidate for demanding professional work while checking the tasks behind each ranking. The scores do not establish broader availability. Artificial Analysis’s September 30 report said access was limited to selected users.

Our Gemini 4 Argon launch report explains the initial rollout and Google’s own benchmark claims. These independent findings add a separate basis for evaluation, with task selection, pricing and the software surrounding the model all affecting the result.

Featured image is Google’s Gemini 4 Argon launch artwork. Image from Google.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile