Cheaper Tokens, Costlier Tasks: Reading the Grok 4.7 Release

xAI released Grok 4.7 on September 21, 2026, and led with a simple pitch: the same price and speed as Grok 4.6, with better performance. The marketing line promised a model twice as fast, at half the price of comparable models.
The performance gains hold up. What the launch framing leaves out is that the price per token and the price per finished task are two different numbers, and only the first one stayed flat. Independent measurement of the second tells a more complicated story about what teams will actually pay.
What xAI Shipped
Grok 4.7 runs on a larger base model than 4.6, trained with a longer reinforcement-learning run weighted toward tasks that take hours to complete. xAI reports gains across coding and knowledge-work benchmarks. On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 reaches 46.3%, compared with 40.4% for Grok 4.6. On DeepSWE v1.1, the published scores rise from 65.2% to 71.0%.
The largest single jump appears on Terminal-Bench 4.0. Grok 4.7 reaches 38.0% versus 20.3% for Grok 4.6, a gain of nearly 18 points. The model also posted strong results in specialized fields, scoring 64.0% on EEBench for electrical engineering, up from 53.0%, and 19.6% on the Harvey Legal Agent Benchmark.
Pricing stayed identical to Grok 4.6: $2 per million input tokens, $6 per million output tokens, with a 500,000-token context window. A separate Grok 4.7 Fast variant runs at roughly double the output speed for double the token price, $4/$12, and is available in Cursor and Grok Build rather than the public API.
Where the "Same Price" Claim Breaks Down
The token rate is not the cost of getting work done. A model priced lower per token can still produce a larger bill if it writes more tokens to reach an answer, and that is exactly what independent testing found.
Independent testing measured how many tokens Grok 4.7 spends per task on a standard intelligence index. Grok 4.7 at its highest reasoning setting used roughly 81,000 output tokens per task, versus 36,000 for Grok 4.6 and 27,000 for the leading OpenAI model. That works out to about 125% more output-token use than Grok 4.6 and nearly triple that of the OpenAI comparison.
That extra output flows straight into cost. Measured at the same reasoning effort rather than xAI's launch configuration, the cost of running one benchmark task rises to $2.73, against Grok 4.6's $1.86, an increase of about 47%. At the top reasoning setting behind most of xAI's launch numbers, the task cost roughly doubles. One test found that an example deck generated with Grok 4.7 at that setting cost about $8, compared with $4.40 for Grok 4.6.
There is a further wrinkle worth flagging for anyone running long prompts. Both models cost $2.00 input, $0.50 cached input and $6.00 output per million tokens below 200,000 tokens of context, and $4.00 / $1.00 / $12.00 at or above it. Cross that context length and the entire request re-bills at the higher rate.
The Reasoning-Effort Problem in the Launch Numbers
The comparison tables xAI published are not a clean model-to-model test. The company reports different reasoning-effort settings for different models, running Grok 4.7 at its highest effort level while showing Grok 4.6 at a lower one, and reporting the Grok 4.7 DeepSWE result at a still different setting.
This matters because the higher effort setting is where much of the extra token consumption comes from. Testing suggests the setting may not even buy better answers. On the intelligence index, both models score the same at high and at the top setting, yet on Grok 4.7 that setting costs $3.74 per task and drops output to under 40 tokens per second. A buyer comparing the launch tables at face value is reading numbers generated at an effort level that raises cost and slows throughput without lifting the score.
The Competitive Picture
On raw capability, Grok 4.7 improves its standing without reaching the top. Independent scoring placed Grok 4.7 at 46 on a widely tracked intelligence index, a two-point gain over Grok 4.6 that puts xAI among the top four AI labs. On CursorBench, it sits between two competitors, outperforming the mid-tier OpenAI model at 41.7% while trailing Anthropic's Claude Fable 5.1 at 51.8%.
The cost-per-task story cuts both ways against Anthropic. Grok 4.7's token consumption raises its own bill relative to its predecessor, but it remains cheaper than the top Claude models on comparable enterprise tasks. On agentic knowledge work, Grok 4.7's per-task cost came in at roughly half that of Claude Opus 5, even as it roughly doubled against Grok 4.6.
The clearest gain shows up in long-horizon agentic work, the category xAI trained toward. On a private multi-hour office-work benchmark, Grok 4.7 gained 111 Elo over its predecessor to reach 1,657, ranking alongside Claude Opus 5 and Claude Fable 5.1 at the frontier.
Safety Claims
xAI paired the release with a new safeguard stack and reported strong results on refusals and jailbreak resistance. The company says the model topped a third-party biosafety benchmark at 62.4%, and on its own internal cyber benchmark, Grok 4.7 blocked all but 3.3% of risky dual-use prompts while rarely blocking legitimate security work. xAI has begun giving select cybersecurity partners invite-only access to the model's red-team capabilities for defense research.
These figures come from xAI's own testing and its internal benchmark, so they carry the usual caveat that applies to vendor-reported safety numbers. They have not yet been independently reproduced.
What It Means for Buyers
The practical takeaway is that the price per token tells you almost nothing about what a workload will cost. Grok 4.7 delivers measurable gains over its predecessor, particularly on terminal work and long-running agentic tasks, and it holds a real price advantage over the top Claude models on enterprise analysis. But the unchanged $2/$6 rate does not mean an unchanged bill. The model writes more to get where it is going, and at the effort setting xAI featured in its launch materials, it can cost roughly twice as much per task as Grok 4.6 for a two-point index gain.
Teams evaluating the upgrade should measure cost per completed task at their own effort setting and prompt length rather than reading the token rate. The two numbers point in different directions, and the gap between them is the whole story of this release.
