xAI โ€“ now branding itself as SpaceXAI on official pages โ€“ shipped Grok 4.7 on September 21, 2026, calling it "our most capable model for coding and knowledge work." After two months of Elon Musk talking about it on X โ€“ 2.1 trillion parameters, SpaceX engineering data, "better than 4.6 in every way" โ€“ the model is finally something you can call through an API. I read the launch post, the docs page, and the independent Artificial Analysis evaluation so you don't have to โ€“ and added the honest context launch posts never give you: what slipped, what the token bill really looks like, and what it means if you run local AI.

This guide explains what Grok 4.7 actually is, which benchmark numbers matter, where it still trails GPT-6 Astra and Claude Fable 5.1, what it costs, and what a new closed flagship means for people running open models on their own hardware.

Source note: specs, prices, and vendor scores below come from xAI's launch announcement, the Grok 4.7 docs page, and Artificial Analysis's independent benchmarking. Musk's pre-release claims are labelled as claims โ€“ several of them were walked back before launch.

The short version

What Grok 4.7 is

Grok 4.7 is xAI's new flagship, successor to Grok 4.6 (shipped August 12). It is a closed, cloud-only model โ€“ no downloadable weights โ€“ with a 500K-token context window, text and image input, text output, and a June 2026 pretraining cutoff with supplemental data through August 2026. The API exposes four reasoning efforts โ€“ low, medium, high (the default), and xhigh โ€“ and several headline benchmarks are reported at xhigh, so read every score with its effort label attached.

Three things went into it, per xAI: an entirely new, larger base model (the 2.1T figure), a longer reinforcement-learning cycle on a harder mix of verifiable tasks skewed toward work that takes many hours, and native training on the Grok Bot harness for conversational and general knowledge work. The SpaceX corpus is the unusual ingredient โ€“ Musk claims it makes 4.7 the best model at "real-world engineering," which is a claim about data uniqueness, not something any public benchmark directly measures.

The road here was bumpy โ€“ and Musk revised the pitch

Worth knowing, because the hype set expectations the launch didn't quite meet. Musk first described 4.7 in July as "better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency," then said on August 12 it would "exceed all current models" and arrive in 3โ€“4 weeks. On September 2 he narrowed it to "10 days" โ€“ pointing at September 12 โ€“ and that date passed without a release.

The delay reason Musk gave is genuinely interesting: the model may have been ending difficult tasks too early and checking its work too weakly because response length was penalized too aggressively in RL โ€“ a "few more days to cook" to fix the reward, not a missing feature. And the claim got quieter: days before launch, Musk reframed 4.7 as "roughly on par with Opus 5.0, better in some ways, worse in others," with multimodal and image handling still needing work โ€“ and pointed past it to Grok 4.8 ("noticeable improvement," early October), 4.9 ("Astra or Fable class"), and Grok 5 ("maybe better than anything," where he expects AGI). No dates or specs for any of those. My read: treat the roadmap as signalling, and judge 4.7 on the shipped numbers below.

The benchmark numbers that matter

First xAI's own comparison table (Grok 4.7 vs 4.6, GPT-5.6 Sol, Fable 5.1), with the effort level xAI reported each score at โ€“ because high vs xhigh is doing real work in this table:

BenchmarkWhat it testsGrok 4.7Grok 4.6Best rival in table
CursorBench 4.0Long-running coding tasks46.3% (xhigh)40.4% (high)Astra-class (see below)
DeepSWE v1.1Repo-scale software engineering71.0% (high)65.2%72.7% (Sol max)
Terminal-Bench 4.0Multi-hour terminal work38.0% (xhigh, Grok Build)20.3% (high)57.9% (Astra)
SWE-Marathon v1.1Marathon coding sessions46.0% (high)31.9%โ€“
EEBenchElectrical-engineering tasks64.0%โ€“โ€“
AA Briefcase v1.1Professional knowledge work1,657 Eloโ€“Opus 5 / Fable 5.1
Harvey Legal AgentLegal agent workflows19.6%โ€“Frontier-class
HealthBench ProfessionalClinical professional tasks56.7% (xhigh)48.5%โ€“

The DeepSWE row is the one that made me sit up: 71.0% at high effort beats Fable 5.1 max (70.0%) and gets close to Sol max (72.7%) โ€“ and it isn't even the xhigh number. Terminal-Bench nearly doubling (20.3% โ†’ 38.0%) is the "works longer" story in one figure. But notice what xAI didn't headline: Terminal-Bench 4.0 at 38.0% is still far behind Astra's 57.9% and Fable 5.1's 55.8%, and the Harvey legal score (19.6%) is modest. This is a strong coding-and-documents release, not a across-the-board takeover.

Now the independent check. Artificial Analysis evaluated Grok 4.7 at xhigh against Grok 4.6 at high on its Intelligence Index: 46 vs 44 โ€“ a +2 nudge that brings SpaceXAI into the top-4 labs but leaves it behind Fable 5.1 and Astra. The gains concentrate exactly where xAI claimed: AA-Briefcase 1,657 Elo (+111, just behind Opus 5 and Fable 5.1, led by analytical quality at 1,994 Elo), GDPval-AA 1,695 (+90), Terminal-Bench 4.0 +4.5pp, GDP.pdf +3.0pp โ€“ with regressions on AA-LCR (โˆ’3.7pp) and AutomationBench-AA (โˆ’1.1pp). On the Coding Agent Index (with Grok Build as harness), it jumps 47 โ†’ 56, ranking 4th behind only Fable 5.1, Astra, and Opus 5, improving on all three components (DeepSWE 65% โ†’ 73%, Terminal-Bench 18% โ†’ 33%, SWE-Atlas-QnA 58% โ†’ 63%). Directionally, the vendor table and the independent rerun agree โ€“ which is more than you can say for most launch weeks.

Safety: the strongest Grok yet, with an invite-only edge

xAI says 4.7 was built with an entirely new safeguard stack, and the numbers are specific enough to take seriously: it tops LatchBio's biosafety benchmark at 62.4% (utility on benign biology tasks plus refusal on dangerous ones), and on HackerBench v0.3 โ€“ risky and malicious cyber tasks โ€“ it lets only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. Hallucination also improved independently: AA-Omniscience hallucination rate 29% vs 34% for 4.6, at roughly flat accuracy (47% vs 48%).

The part to watch: xAI has started giving select cybersecurity partners invite-only access to 4.7's red-team capabilities for defense research. That is the same playbook as every frontier lab now โ€“ ship the safe model publicly, gate the sharp edges privately. If you run anything exposed to the internet, the practical read is the same as Astra's: patch windows are getting shorter in both directions.

Price and availability โ€“ the honest math

The sticker price is unchanged from Grok 4.6, and it is genuinely cheap for a flagship:

Now the part launch posts skip: throughput cost. Artificial Analysis measured ~81K output tokens per Intelligence Index task for 4.7 at xhigh, vs ~36โ€“38K for 4.6 and ~27K for Astra at max โ€“ 125โ€“196% more. At those volumes the estimated cost per task is ~$3.74 for Grok 4.7 vs ~$3.26 for Astra despite the 5x sticker gap. Output speed measured ~188 tokens/second, ~7.1 minutes per task. So "twice as fast, at half the price of comparable models" is true per token and roughly true per minute โ€“ but per finished task, the token appetite eats the discount. If you build agents on this, budget on tasks, not tokens โ€“ our AI API cost calculator does exactly that math with your own mix.

What this means if you run local AI

The question readers of this site actually have: does Grok 4.7 change anything for my RX 6800M? Three precise answers.

Where the gap widened: long-horizon agentic coding. DeepSWE 71%, SWE-Marathon 46%, CursorBench 46.3% โ€“ these are "start Friday, review Monday" workloads no 27Bโ€“35B model on 12 GB of VRAM touches. If your workflow is delegating multi-hour repo work, cloud just pulled further ahead, and Grok Build being free to try makes it cheap to verify on your own repos.

Where nothing changed: privacy, offline, and per-task cost at small scale. A local Qwen or Tiel-Coder MoE costs electricity and keeps every byte on your machine; Grok needs your code on someone else's servers. And a short, well-scoped task that a local 32B model finishes in one pass is still effectively free vs a billed API call. Our lab benchmarks show what local models are actually good at: fast, private, repeatable generation.

The interesting bit: Grok 4.7's delay story is a free lesson in RL reward design โ€“ penalize length too hard and the model quits early and stops checking its work. If you fine-tune or RL your own small models, that failure mode ("ends difficult tasks too early") is worth checking for before you blame the base model. Frontier labs' mistakes are open-source education.

Bottom line

Grok 4.7 is a real, well-measured step โ€“ the best Grok ever on long coding and professional knowledge work, with genuinely strong safety numbers โ€“ but it is not the "exceeds all current models" release Musk previewed in August. It is roughly Opus-5-class, 4th on the independent coding index, still well behind Astra and Fable 5.1 on terminal work, and its cheap tokens come with a heavy token habit. Judge it per finished task, not per million tokens โ€“ and keep an eye on whether 4.8 in October is the one that actually closes the gap.

Sources