GPT 6.1 Sol joins Agent Arena with a lower median task cost

Arena’s October 2 addition puts GPT 6.1 Sol alongside its predecessor, with a lower observed median task cost and uncertainty that matters for interpretation.
Accessibility Adjustments
Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

Arena’s October 2 addition puts GPT 6.1 Sol alongside its predecessor, with a lower observed median task cost and uncertainty that matters for interpretation.

Arena details a training recipe that combines human preferences with checks on image instructions. Its new disclosure uses older leaderboard results.

New Agent Arena results place Sonnet 5.5 near the top while showing why effort settings and measured task spending matter alongside token prices.

Original October 1 screenshots show Grok 4.7 labels in two consumer app modes, while plan eligibility and wider access remain unconfirmed.

Ai2 opens its scientific report model and training data, with faster generation, older evaluation baselines and different licenses for weights and data.

OpenAI expands its personal finance tool to more U.S. users, with connected account insights but no ability to move money or make trades.

New research reports a citation decline in retail and travel prompts without establishing why.

ServiceNow reports improved enterprise task performance using generated training examples.

Boston Dynamics reveals an Atlas hand designed for manipulation, simulation and the demands of industrial work.

New research examines a more selective approach to screening biological AI work.

Eni and Generative Bionics will assess humanoid use cases, robot materials and computing support under a new research and industrial agreement.

Honda is evaluating a robotic hand and Redwire arm combination to handle laboratory tasks aboard future commercial space stations.

The October policy outlines responsibilities for researchers and analysts using AI assisted work.

A new cap changes how researchers schedule papers and share responsibility for AI assisted work.

Google Cloud adds Grok 4.7 in preview with shared endpoint quotas, a large context window and separate pricing for long inputs.

Perplexity Decider 27B combines public model weights with a hosted API that returns probabilities for classification and routing tasks.

Strands Decider 2B selects options and scores inputs for AI agent workflows, with open weights and important limits on reasoning and confidence.

Google Cloud makes its agent topology API generally available with a custom query builder. Supported resources, security data limits and upcoming billing matter.

Google finds a different mix of vulnerabilities in AI attributed research, with important limits on what the comparison proves.

Suno Speech creates spoken narration and background music in one track. The public beta is available on web and mobile, with uneven accents and timing still possible.

Black Forest Labs opens a dedicated FLUX 3 image endpoint with layout controls, reference editing and output up to 4K. Pricing and preservation limits matter.

Matthew Schwartz releases a toolkit built around checkable calculations and human direction

September releases improve provider restrictions model fallback and gateway usage accounting

September 14 screenshots show a Claude Money tab and bank linking prompt. A public launch, supported banks and access permissions remain unconfirmed.