Grok 4.7 is cheap per token and mid-pack on agentic coding
SpaceXAI prices its new coding model at half the comparable rate. Independent scoring puts it well behind Claude and GPT-6 on the agentic tasks it targets.
The vendor page and the independent index disagree about the same model. Both numbers are below.
SpaceXAI, the company formed when SpaceX acquired xAI in February 2026, released Grok 4.7 on September 21. The pitch on its own page is twice as fast at half the price of comparable models, at $2 per million input tokens and $6 per million output tokens, with a Fast variant at double the price for double the output speed.
SpaceXAI's own numbers: CursorBench 4.0 at 46.3 percent, up from 40.4 for Grok 4.6. DeepSWE v1.1 at 71.0 percent on high effort. EEBench, on electrical engineering, at 64.0 percent. On safety, the company reports its HackerBench v0.3 run allows 3.3 percent of risky dual-use prompts, and describes an entirely new safeguard stack. Availability is the Grok API, Cursor, and Grok Build.
The independent column tells a different story
The Decoder ran the comparison against the Artificial Analysis Intelligence Index v4.3.2 and put Grok 4.7 at 46, with Claude Fable 5.1 and GPT-6 both at 53. On Terminal-Bench 4.0, which measures agentic coding rather than single-turn answers, the spread is wider: Grok 4.7 at 26 percent, against 60 for GPT-6 Astra, 55 for Claude Fable 5.1, and 27 for DeepSeek V4.1 Flash. The Decoder's summary is that the model lands mid-pack overall and that the gap grows wider in agentic coding.
Both sets of numbers can be true at once, because they measure different things. CursorBench and DeepSWE are the benchmarks SpaceXAI chose to publish. Terminal-Bench is the one that most resembles an agent left alone with a shell for an hour, which is the job the price is trying to win.
What nobody has published is a cost-per-completed-task comparison, which is the only figure that would settle whether cheap tokens beat a higher success rate. A model at a third of the price that needs three attempts costs the same and takes three times as long.
Why a build studio cares
We put models behind a router, and a router entry is a per-task decision rather than a favorite. At $2 in and $6 out, Grok 4.7 is priced into the bracket where we would use it for high-volume, low-stakes, single-pass work: classification, summarizing a page, drafting a first cut somebody will edit. A 26 percent Terminal-Bench score is the number that keeps it out of the unattended slot, where a failed run costs a rollback rather than a retry. The thing we will not do is treat a launch-day price as a standing reason to move traffic, because launch-day prices move: in the last two months we have covered an 80 percent cut to GPT-5.6 Luna and DeepSeek cutting its Flash pricing.
Next step: read SpaceXAI's release notes next to The Decoder's comparison, and work out your own cost per completed task before switching anything. If you want help building that measurement, write to us at hello@gattyworks.com.