# Grok 4.7: xAI's 2.1T SpaceX-trained flagship, explained

> xAI shipped Grok 4.7 on Sept 21, 2026: 2.1T params, SpaceX-data training, 500K context, $2/$6 pricing, DeepSWE 71% and AA Coding Index 56. What improved, what didn't, and what it means for local AI.

*Source: https://velstech.net/grok-4-7 · Updated: 2026-09-22 · Category: AI · Tags: LLM, xAI, Grok, Benchmarks, AI News*

*Markdown version of [Grok 4.7: xAI's 2.1T SpaceX-trained flagship, explained](https://velstech.net/grok-4-7). [Read the full guide with interactive tools](https://velstech.net/grok-4-7).*
*Also as Markdown: [Hindi](https://velstech.net/grok-4-7.hi.md) · [Tamil](https://velstech.net/grok-4-7.ta.md).*

---

xAI – now branding itself as **SpaceXAI** on official pages – shipped
**Grok 4.7** on September 21, 2026, calling it "our most capable model
for coding and knowledge work." After two months of Elon Musk talking about it on X –
2.1 trillion parameters, SpaceX engineering data, "better than 4.6 in every way" –
the model is finally something you can call through an API. I read the launch post,
the docs page, and the independent Artificial Analysis evaluation so you don't have
to – and added the honest context launch posts never give you: what slipped, what
the token bill really looks like, and what it means if you run local AI.

This guide explains what Grok 4.7 actually is, which benchmark numbers matter, where
it still trails [GPT-6 Astra](https://velstech.net/gpt-6-astra) and
[Claude Fable 5.1](https://velstech.net/claude-fable-5-1), what it costs, and what a new
closed flagship means for people running open models on their own hardware.

*Source note:* specs, prices, and vendor scores below come from xAI's
[launch announcement](https://x.ai/news/grok-4-7), the
[Grok 4.7 docs page](https://docs.x.ai/developers/grok-4-7), and
[Artificial Analysis's
independent benchmarking](https://artificialanalysis.ai/articles/benchmarking-grok-4-7). Musk's pre-release claims are labelled as claims –
several of them were walked back before launch.

## The short version

- New larger pretrain: ~2.1T parameters per Musk (unverified by any
model card), up from the 1.5T base behind Grok 4.6 – plus supplemental training on
a large SpaceX engineering corpus and a longer RL cycle weighted toward
multi-hour tasks.

- Longer horizons are the headline. CursorBench 4.0 up 40.4% → 46.3%,
DeepSWE v1.1 up 65.2% → 71.0% (past Fable 5.1, near GPT-5.6 Sol), Terminal-Bench 4.0
up 20.3% → 38.0%. It works longer on hard tasks and checks its work more carefully.

- Independent verdict: better agent, same neighbourhood. Artificial
Analysis scores it 46 on the Intelligence Index (+2 over 4.6) and 56 on the Coding
Agent Index (+9) – 4th overall, behind Fable 5.1, Astra, and Opus 5.

- The catch is token appetite. At max reasoning effort it burns ~81K
output tokens per task – 2–3x its rivals – so the $2/$6 sticker price works out to
roughly the same cost per finished task as the $10/$50 flagships.

- Same price as 4.6: $2 per million input / $6 output under 200K
tokens (doubling above that), cached input $0.50, with a 2x-speed fast variant at
2x price. Live in Cursor, Grok Build (free to try), the Grok API as
grok-4-7, routers, and clouds.

## What Grok 4.7 is

Grok 4.7 is xAI's new flagship, successor to Grok 4.6 (shipped August 12). It is a
closed, cloud-only model – no downloadable weights – with a 500K-token
[context window](https://velstech.net/what-is-an-llm), text and image input, text output,
and a June 2026 pretraining cutoff with supplemental data through August 2026. The
API exposes four reasoning efforts – low, medium, high (the default), and xhigh –
and several headline benchmarks are reported at xhigh, so read every score with its
effort label attached.

Three things went into it, per xAI: an entirely new, larger base model (the 2.1T
figure), a longer reinforcement-learning cycle on a harder mix of verifiable tasks
skewed toward work that takes many hours, and native training on the Grok Bot
harness for conversational and general knowledge work. The SpaceX corpus is the
unusual ingredient – Musk claims it makes 4.7 the best model at "real-world
engineering," which is a claim about data uniqueness, not something any public
benchmark directly measures.

## The road here was bumpy – and Musk revised the pitch

Worth knowing, because the hype set expectations the launch didn't quite meet. Musk
first described 4.7 in July as "better than 4.6 in every way, except slightly slower
to serve, albeit with even better token efficiency," then said on August 12 it would
"exceed all current models" and arrive in 3–4 weeks. On September 2 he narrowed it
to "10 days" – pointing at September 12 – and that date passed without a release.

The delay reason Musk gave is genuinely interesting: the model may have been ending
difficult tasks too early and checking its work too weakly because response length
was penalized too aggressively in RL – a "few more days to cook" to fix the
reward, not a missing feature. And the claim got quieter: days before launch, Musk
reframed 4.7 as "roughly on par with Opus 5.0, better in some ways, worse in
others," with multimodal and image handling still needing work – and pointed past
it to Grok 4.8 ("noticeable improvement," early October), 4.9 ("Astra or Fable
class"), and Grok 5 ("maybe better than anything," where he expects AGI). No dates
or specs for any of those. My read: treat the roadmap as signalling, and judge 4.7
on the shipped numbers below.

## The benchmark numbers that matter

First xAI's own comparison table (Grok 4.7 vs 4.6, GPT-5.6 Sol, Fable 5.1), with the
effort level xAI reported each score at – because high vs xhigh is doing real work
in this table:

| Benchmark | What it tests | Grok 4.7 | Grok 4.6 | Best rival in table |
| --- | --- | --- | --- | --- |
| CursorBench 4.0 | Long-running coding tasks | 46.3% (xhigh) | 40.4% (high) | Astra-class (see below) |
| DeepSWE v1.1 | Repo-scale software engineering | 71.0% (high) | 65.2% | 72.7% (Sol max) |
| Terminal-Bench 4.0 | Multi-hour terminal work | 38.0% (xhigh, Grok Build) | 20.3% (high) | 57.9% (Astra) |
| SWE-Marathon v1.1 | Marathon coding sessions | 46.0% (high) | 31.9% | – |
| EEBench | Electrical-engineering tasks | 64.0% | – | – |
| AA Briefcase v1.1 | Professional knowledge work | 1,657 Elo | – | Opus 5 / Fable 5.1 |
| Harvey Legal Agent | Legal agent workflows | 19.6% | – | Frontier-class |
| HealthBench Professional | Clinical professional tasks | 56.7% (xhigh) | 48.5% | – |

The DeepSWE row is the one that made me sit up: 71.0% at *high* effort beats
Fable 5.1 max (70.0%) and gets close to Sol max (72.7%) – and it isn't even the
xhigh number. Terminal-Bench nearly doubling (20.3% → 38.0%) is the "works longer"
story in one figure. But notice what xAI didn't headline: Terminal-Bench 4.0 at
38.0% is still far behind Astra's 57.9% and Fable 5.1's 55.8%, and the Harvey legal
score (19.6%) is modest. This is a strong coding-and-documents release, not a
across-the-board takeover.

Now the independent check. Artificial Analysis evaluated Grok 4.7 at xhigh against
Grok 4.6 at high on its Intelligence Index: **46 vs 44** – a +2 nudge
that brings SpaceXAI into the top-4 labs but leaves it behind Fable 5.1 and Astra.
The gains concentrate exactly where xAI claimed: AA-Briefcase 1,657 Elo (+111,
just behind Opus 5 and Fable 5.1, led by analytical quality at 1,994 Elo),
GDPval-AA 1,695 (+90), Terminal-Bench 4.0 +4.5pp, GDP.pdf +3.0pp – with regressions
on AA-LCR (−3.7pp) and AutomationBench-AA (−1.1pp). On the Coding Agent Index (with
Grok Build as harness), it jumps 47 → **56**, ranking 4th behind only
Fable 5.1, Astra, and Opus 5, improving on all three components (DeepSWE 65% → 73%,
Terminal-Bench 18% → 33%, SWE-Atlas-QnA 58% → 63%). Directionally, the vendor table
and the independent rerun agree – which is more than you can say for most launch
weeks.

## Safety: the strongest Grok yet, with an invite-only edge

xAI says 4.7 was built with an entirely new safeguard stack, and the numbers are
specific enough to take seriously: it tops LatchBio's biosafety benchmark at 62.4%
(utility on benign biology tasks plus refusal on dangerous ones), and on HackerBench
v0.3 – risky and malicious cyber tasks – it lets only 3.3% of risky dual-use prompts
through while rarely blocking legitimate security work. Hallucination also improved
independently: AA-Omniscience hallucination rate 29% vs 34% for 4.6, at roughly flat
accuracy (47% vs 48%).

The part to watch: xAI has started giving select cybersecurity partners invite-only
access to 4.7's red-team capabilities for defense research. That is the same
playbook as every frontier lab now – ship the safe model publicly, gate the sharp
edges privately. If you run anything exposed to the internet, the practical read is
the same as Astra's: patch windows are getting shorter in both directions.

## Price and availability – the honest math

The sticker price is unchanged from Grok 4.6, and it is genuinely cheap for a
flagship:

- API: $2 per million input / $6 output under 200K prompt tokens; at or above 200K the whole request bills at $4/$12. Cached input $0.50. A fast variant gives 2x output speed at 2x price.

- Try it free: Grok Build offers free access; also live in Cursor, third-party coding harnesses, model routers, and cloud platforms. API model ID grok-4-7.

- Context: 500K tokens, unchanged from 4.6.

Now the part launch posts skip: throughput cost. Artificial Analysis measured ~81K
output tokens per Intelligence Index task for 4.7 at xhigh, vs ~36–38K for 4.6 and
~27K for Astra at max – 125–196% more. At those volumes the estimated cost per task
is ~$3.74 for Grok 4.7 vs ~$3.26 for Astra *despite* the 5x sticker gap.
Output speed measured ~188 tokens/second, ~7.1 minutes per task. So "twice as fast,
at half the price of comparable models" is true per token and roughly true per
minute – but per *finished task*, the token appetite eats the discount. If
you build agents on this, budget on tasks, not tokens – our
[AI API cost calculator](https://velstech.net/ai-api-cost-calculator) does exactly that
math with your own mix.

## What this means if you run local AI

The question readers of this site actually have: does Grok 4.7 change anything for
my RX 6800M? Three precise answers.

**Where the gap widened:** long-horizon agentic coding. DeepSWE 71%,
SWE-Marathon 46%, CursorBench 46.3% – these are "start Friday, review Monday"
workloads no 27B–35B model on 12 GB of VRAM touches. If your workflow is delegating
multi-hour repo work, cloud just pulled further ahead, and Grok Build being free to
try makes it cheap to verify on your own repos.

**Where nothing changed:** privacy, offline, and per-task cost at
small scale. A [local Qwen or Tiel-Coder MoE](https://velstech.net/llama-cpp-guide) costs electricity and keeps every byte
on your machine; Grok needs your code on someone else's servers. And a short,
well-scoped task that a local 32B model finishes in one pass is still effectively
free vs a billed API call. Our [lab benchmarks](https://velstech.net/benchmarks/index)
show what local models are actually good at: fast, private, repeatable generation.

**The interesting bit:** Grok 4.7's delay story is a free lesson in RL
reward design – penalize length too hard and the model quits early and stops
checking its work. If you fine-tune or RL your own small models, that failure mode
("ends difficult tasks too early") is worth checking for before you blame the base
model. Frontier labs' mistakes are open-source education.

## Bottom line

Grok 4.7 is a real, well-measured step – the best Grok ever on long coding and
professional knowledge work, with genuinely strong safety numbers – but it is not
the "exceeds all current models" release Musk previewed in August. It is roughly
Opus-5-class, 4th on the independent coding index, still well behind Astra and
Fable 5.1 on terminal work, and its cheap tokens come with a heavy token habit.
Judge it per finished task, not per million tokens – and keep an eye on whether
4.8 in October is the one that actually closes the gap.

## Sources

- xAI / SpaceXAI: Introducing Grok 4.7 (Sept 21, 2026)

- SpaceXAI Docs: Grok 4.7 technical overview – 500K context, pricing, reasoning effort

- Artificial Analysis: Benchmarking Grok 4.7 – Intelligence Index 46, Coding Agent Index 56

- heise online: Grok 4.7 – inexpensive API meets high token consumption

- TestingCatalog: SpaceXAI releases Grok 4.7 for coding and knowledge work

## FAQ

**What is Grok 4.7?**

Grok 4.7 is xAI's (SpaceXAI's) flagship model for coding and knowledge work, released Sept 21, 2026 – a ~2.1T-parameter pretrain with SpaceX-data supplemental training, 500K context, and reasoning efforts from low to xhigh via API model grok-4-7.

**How much does Grok 4.7 cost?**

$2 per million input and $6 output under 200K prompt tokens (whole request doubles to $4/$12 above that), $0.50 cached input, with a 2x-speed fast variant at 2x price. But at ~81K output tokens per task, per-task cost roughly matches $10/$50 flagships.

**Is Grok 4.7 better than GPT-6 Astra or Claude Fable 5.1?**

Not overall. It beats Fable 5.1 on DeepSWE (71.0% vs 70.0%) and ranks 4th on the independent Coding Agent Index, but trails Astra and Fable 5.1 on Terminal-Bench 4.0 and the overall Intelligence Index (46 vs 61+). Musk himself reframed it as roughly Opus-5-class.

**Can I run Grok 4.7 locally?**

No. It is a closed, cloud-only model available in Cursor, Grok Build (free to try), the Grok API, routers, and clouds. You cannot download its weights or run it on your own GPU.

---

*VelsTech – technology explained for everyone. Original: https://velstech.net/grok-4-7*
