Grok 4.7 reaches Google Cloud Model Garden in preview
Google Cloud adds Grok 4.7 in preview with shared endpoint quotas, a large context window and separate pricing for long inputs.

Google Cloud added Grok 4.7 to Model Garden in Preview on September 30, according to its Gemini Enterprise Agent Platform release notes. The addition gives developers another managed route to use xAIโs model for coding and knowledge work.
xAIโs original model announcement was published on September 21, when Grok 4.7 was already available through its API, Cursor and Grok Build. Googleโs September 30 release adds a separate cloud deployment option with its own availability and usage limits.
What the preview supports
Googleโs dedicated model page lists text and image inputs with text output. Function calling, structured output and reasoning are supported in Preview. Batch predictions are unavailable. The model has a context window of 524,288 tokens.
Access uses fixed quotas. Google lists its Standard pay as you go mode and Provisioned Throughput as unsupported for this preview. Availability includes the global endpoint and the US multiregion endpoint, while the model page lists machine learning processing in the United States.
Two endpoints share one allowance
The published limits are 13 queries per minute, 188,000 input tokens per minute and 16,000 output tokens per minute. These are separate request and token ceilings, so applications need to account for both when planning workloads.
The Grok integration guide says the global and US endpoints draw from the same quota for each base model. Switching endpoints therefore does not create another allowance. Google also warns that account limits can vary and access can be restricted. Developers should inspect their projectโs actual quotas before estimating capacity.
Long inputs change the token price
Googleโs pricing table lists Grok 4.7 at $2 per million input tokens and $6 per million output tokens for requests with no more than 200,000 input tokens. Cache hits cost $0.50 per million tokens. All prices are in US dollars.
Above 200,000 input tokens, the rates rise to $4 for input, $12 for output and $1 for cache hits. Google says the higher rates apply to all tokens in that request. A larger context window therefore carries both capacity and billing considerations.
A deployment option with preview limits
Google labels the offering Preview, with potentially limited support. Teams evaluating it should check quota availability, endpoint requirements and their own tool workflows before moving a production workload. ByteForward has not independently tested this integration.
Related coverage explains Googleโs Cloud Trace integration for agent requests and App Topology tools for mapping agent dependencies.
Featured image is an original AI generated conceptual editorial illustration.



