xAI's API release notes list a new frontier model in August 2026, Grok 4.6, and another in September, Grok 4.7. Both carry the same description, a model "for coding, agentic tasks, and knowledge work", and most of the same specs.
- context window, 4.6 and 4.7
- 500k
- input / output per 1M tokens, under 200k
- $2 / $6
- reasoning effort levels, low to xhigh
- 4
Shared specs
Both take text and image input, return text, and accept a reasoning effort of low, medium, high or xhigh. Both use a two-tier price that depends on prompt length:
| Prompt size | Input | Cached input | Output |
|---|---|---|---|
| under 200k tokens | $2 | $0.50 | $6 |
| over 200k tokens | $4 | $1 | $12 |
Design around the jump at 200k. With a 500k window you can paste a whole repository into one request, and past 200k every token in that request costs twice as much. For code work I retrieve the files a task needs and stay under the line. Cached input costs a quarter of fresh input, so keep stable context first and unchanged between calls.
New in 4.7
The release notes give 4.7 three things 4.6 doesn't have:
- A US regional endpoint, for teams with data-residency rules.
- A "Grok 4.7 Fast" variant, available only inside Cursor and xAI's Grok Build tool, not on the API.
- On the Responses API,
grok-4.7always returnsreasoning.encrypted_content, even whenincludedoesn't ask for it.
The reasoning effort setting deserves a test on either model. For a code-review bot, low on small diffs and high on large ones keeps latency and cost in proportion to the change. The parameter defaults to high, one step below xhigh, so a bot that leaves it out pays for more reasoning than a small diff needs.
const response = await fetch("https://api.x.ai/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.XAI_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "grok-4.7",
reasoning_effort: diffLines > 400 ? "high" : "low",
messages: [{ role: "user", content: reviewPrompt }],
}),
});
Around the models
- Grok Bot (August): durable AI teammates that run on a persistent cloud computer, with messaging and connectors.
- Imagine: an
autoquality default, up to five source images for edits, and new 21:9 and 5:2 aspect ratios. - Voice: Grok Voice Transcribe 2.0 for speech-to-text. Version 1.0 stays the default.
- Priority processing (June): a
service_tier: "priority"option on text endpoints.
On the calendar: November 2
The grok-imagine-image-quality model retires on November 2, 2026. Its requests will route to grok-imagine-image-2.0 with quality set to low, at a lower per-image price. If your product depends on the old model's output quality, pin the new model and a quality setting yourself before that date.
The release notes list no Grok 5 yet. With two releases a month apart at one price, I put the model id in configuration so the next switch is a one-line change.
Sources
- xAI API release notes, June to September 2026
- xAI releases Grok 4.6 flagship model, DataNorth AI
Volodymyr Chornous

