Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Grok 4.7 reaches Google Cloud Model Garden in preview

Google Cloud adds Grok 4.7 in preview with shared endpoint quotas, a large context window and separate pricing for long inputs.

Listen to this article

Google Cloud added Grok 4.7 to Model Garden in Preview on September 30, according to its Gemini Enterprise Agent Platform release notes. The addition gives developers another managed route to use xAI’s model for coding and knowledge work.

xAI’s original model announcement was published on September 21, when Grok 4.7 was already available through its API, Cursor and Grok Build. Google’s September 30 release adds a separate cloud deployment option with its own availability and usage limits.

What the preview supports

Google’s dedicated model page lists text and image inputs with text output. Function calling, structured output and reasoning are supported in Preview. Batch predictions are unavailable. The model has a context window of 524,288 tokens.

Access uses fixed quotas. Google lists its Standard pay as you go mode and Provisioned Throughput as unsupported for this preview. Availability includes the global endpoint and the US multiregion endpoint, while the model page lists machine learning processing in the United States.

Two endpoints share one allowance

The published limits are 13 queries per minute, 188,000 input tokens per minute and 16,000 output tokens per minute. These are separate request and token ceilings, so applications need to account for both when planning workloads.

The Grok integration guide says the global and US endpoints draw from the same quota for each base model. Switching endpoints therefore does not create another allowance. Google also warns that account limits can vary and access can be restricted. Developers should inspect their project’s actual quotas before estimating capacity.

Long inputs change the token price

Google’s pricing table lists Grok 4.7 at $2 per million input tokens and $6 per million output tokens for requests with no more than 200,000 input tokens. Cache hits cost $0.50 per million tokens. All prices are in US dollars.

Above 200,000 input tokens, the rates rise to $4 for input, $12 for output and $1 for cache hits. Google says the higher rates apply to all tokens in that request. A larger context window therefore carries both capacity and billing considerations.

A deployment option with preview limits

Google labels the offering Preview, with potentially limited support. Teams evaluating it should check quota availability, endpoint requirements and their own tool workflows before moving a production workload. ByteForward has not independently tested this integration.

Related coverage explains Google’s Cloud Trace integration for agent requests and App Topology tools for mapping agent dependencies.

Featured image is an original AI generated conceptual editorial illustration.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile