xAI introduces Grok 4.5 for coding and agent tasks
The July 16 Grok 4.5 announcement details launch pricing and reported coding results. This update corrects the earlier date and separates benchmarks from deployment decisions.

Correction added October 1, 2026. The official Grok 4.5 announcement is dated July 16, rather than July 8. It introduced a model for coding, agent tasks, and knowledge work, with launch pricing of $2 per million input tokens and $6 per million output tokens.
The launch pairs a coding model with a clear token price. The practical question is how much useful work a team gets for its total spend.
Pricing and cost per completed task
A low token price does not establish a low cost per successful task. Retries, the length of model output, and any supporting tools can change the final bill.
Teams comparing products should use the same workload and account for the complete workflow. Model access through an API and a bundled coding product can have different costs and capabilities.
Coding benchmark claims
The launch page reports 83.3% for Grok 4.5 on Terminal Bench 2.1, alongside 83.4% for GPT 5.5 in the published comparison. These are results presented by the developer, not an independent ByteForward evaluation.
The comparison draws competitor figures from their published system cards or benchmark leaderboards. Similar scores on one test do not establish equivalent behavior across every engineering task.
API availability
The announcement lists Grok Build, Cursor, and the SpaceXAI API as launch access routes. Available tools and usage limits depend on the product through which the model is used.
Before connecting a coding tool to a project, review its permissions, retention terms, and the actions it can take. A model benchmark is separate from those operational choices.
Who Grok 4.5 is for
Grok 4.5 is relevant to teams comparing models for coding and agent work. For the later release, see our coverage of Grok 4.6. The two launches should be evaluated using their own dated evidence.
Readers comparing coding models should evaluate successful task completion and total cost on their own workloads, rather than infer equivalence from a single benchmark.



