Grok 4.6 has 500,000 tokens — but Azure documents 200,000

The same model name appears on two proposals. One promises a 500,000-token context window. The other sets the limit at 200,000. Both are called Grok 4.6.
That is not a footnote for developers. It decides which documents fit, how much room the model has to answer, and whether the solution you sold can run on the selected cloud platform.
The same model name, two different products
SpaceXAI describes Grok 4.6 with 500,000 tokens in both its Amazon Bedrock and Microsoft Foundry announcements. Microsoft's own implementation guide, updated on August 26, instead documents a total 200,000-token context window for grok-4.6 version 1.
A cloud route is the endpoint and commercial delivery through which a model is used. The route can change context, tools, quotas, geography, support terms, and billing while the model name stays the same.
This does not automatically make either side wrong if the underlying model supports more than a particular endpoint exposes. For the buyer, however, Microsoft's documented endpoint is the contract that can be planned against until the documentation or service changes.
Source: SpaceXAI — Grok 4.6 on Microsoft Foundry and Microsoft — Deploy and use Grok models in Microsoft Foundry.
The 300,000-token gap is not cosmetic
Consider a 190,000-token input: agreements, appendices, history, and instructions. In a 500,000-token context window, it occupies 38%. In Foundry's documented window, the same input occupies 95%.
Microsoft also says that input and generated output share the 200,000-token budget. A 190,000-token input therefore leaves no more than 10,000 tokens for visible output and reasoning, even though the stated maximum output is 128,000. A job that fits through a 500,000-token route may need splitting, compression, or architectural changes on Foundry.
This is where model comparisons often fail. A benchmark can say something about quality. It says nothing about whether your combination of source material, tools, and output fits through the endpoint you intend to buy.
Source: Microsoft — token limits and context window for Grok 4.6 and SpaceXAI — Grok 4.6 on Amazon Bedrock.
Matching list prices hide different deliveries
On the surface, the pricing is aligned. SpaceXAI lists Bedrock at $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. Microsoft's launch lists the same three rates for Global Standard.
Matching unit prices do not make the routes economically equivalent. If the 200,000-token limit requires more calls, aggressive compression, or a separate retrieval layer, both development cost and failure modes change. If a larger route can take the entire source in one request, it may cost less to operate despite the same token price. The reverse can also be true if the larger window encourages wasteful prompts.
The useful measure is not price per token in a launch announcement. It is cost per accepted outcome in the target environment, including extra calls, cache hits, tool runs, and rework.
Source: Microsoft — Grok 4.6 comes to Microsoft Foundry Models and SpaceXAI — Bedrock pricing.
The cloud route determines more than context
The Foundry offer is in Public Preview without a preview SLA. Microsoft documents only Global Standard for Grok 4.6, which means processing may occur globally; Data Zone Standard is not available for this version at publication time. The route supports Chat Completions and Responses, text and image input, function calling, streaming, and JSON. Microsoft also notes that some SpaceXAI options, including live search, do not apply to Azure-hosted deployments.
The Bedrock launch, by contrast, is marked generally available in supported AWS Regions and advertises 500,000 tokens. That does not prove that every AWS account, region, tool, or guardrail behaves identically. Those details have to be read back in the region and account that will carry the customer solution.
Two cloud routes can therefore answer four business questions differently: Does the workload fit? Where is it processed? Which capabilities exist? What operating promise can you make to the customer?
Source: Microsoft's Foundry guide and SpaceXAI's Bedrock launch.
The customer proposal should name the endpoint
“We use Grok 4.6” is no longer a sufficient architecture description. It leaves open whether you mean the SpaceXAI API, Bedrock, Foundry, or another partner — and therefore which context, geography, quota, and support level actually applies.
A proposal is more honest when it names the model version, provider route, deployment type, and limits verified in the customer's account. It can then price change explicitly: what happens if a preview becomes GA, a regional option appears, or the context limit increases?
This is a procurement question as much as a technical one. Anyone comparing only model names risks buying a capability that exists in the marketing but not in the endpoint the organization is allowed to use.
Buy evidence from the route you will operate
The documentation may converge after this article is published. That does not change the conclusion. Before pricing a customer solution, run the same representative workloads through the intended endpoint with the actual model version, region, tools, and quota. Capture accepted input size, available output, tool behavior, latency, and the invoice line.
If the results differ, the model has not “failed.” You have discovered that the distribution route is part of the product.
When cloud route, cost, and endpoint behavior must be compared before a proposal, Tool Forge can help build and measure a bounded evaluation. The goal is not to declare a universal winner, but to select the route that can actually keep your promise to the customer.
FAQ
Does Grok 4.6 have a 200,000- or 500,000-token context window?
It depends on the distribution route. Microsoft documents a total 200,000-token window for Grok 4.6 in Foundry, while SpaceXAI describes 500,000 tokens in both its Bedrock and Foundry announcements. Plan against the contract and behavior of the endpoint you actually buy.
Is Grok 4.6 production-ready in Microsoft Foundry?
Microsoft classifies the route as Public Preview without a preview SLA and documents Global Standard as the only deployment type at publication time. Evaluate the endpoint against your production requirements before making customer commitments.
What should buyers compare when the same AI model is offered by several clouds?
Compare the exact model version, total context and output limits, tool support, data geography, preview or GA status, SLA, quotas, cache pricing, and the actual invoice in the target environment. The model name is not enough.
The Forge newsletter
Get new articles in your inbox
Pick the topics you care about. No noise, at most one email a week.
We follow GDPR. Unsubscribe anytime.


