xAI – जो अब official pages पर खुद को SpaceXAI कहता है – ने 21 सितंबर 2026 को Grok 4.7 लॉन्च किया, इसे "coding और knowledge work के लिए हमारा सबसे capable model" बताया। X पर Elon Musk के दो महीने के दावों के बाद – 2.1 trillion parameters, SpaceX engineering डेटा, "हर तरह से 4.6 से बेहतर" – यह model अब आखिरकार API से इस्तेमाल किया जा सकता है। मैंने launch post, docs page और Artificial Analysis का independent evaluation पढ़ा है ताकि आपको न पढ़ना पड़े – और वह ईमानदार context भी जोड़ा है जो launch posts कभी नहीं देते: क्या टला, token का असली बिल कितना है, और अगर आप local AI चलाते हैं तो इसका क्या मतलब है।
यह guide बताती है कि Grok 4.7 असल में क्या है, कौन-से benchmark नंबर मायने रखते हैं, कहाँ यह अब भी GPT-6 Astra और Claude Fable 5.1 से पीछे है, इसकी कीमत क्या है, और अपने hardware पर open models चलाने वालों के लिए एक नए closed flagship का क्या मतलब है।
Source note: नीचे specs, कीमतें और vendor scores xAI के launch announcement, Grok 4.7 docs page और Artificial Analysis के independent benchmarking से हैं। Musk के pre-release दावे जहाँ-तहाँ "दावे" कहकर लेबल किए गए हैं – उनमें से कई launch से पहले वापस ले लिए गए थे।
संक्षेप में
- नया बड़ा pretrain: Musk के अनुसार ~2.1T parameters (किसी model card से unverified), Grok 4.6 के 1.5T base से ज़्यादा – साथ में SpaceX engineering corpus पर supplemental training और multi-hour tasks पर ज़ोर वाला लंबा RL cycle।
- लंबे horizons ही headline हैं। CursorBench 4.0 40.4% → 46.3%, DeepSWE v1.1 65.2% → 71.0% (Fable 5.1 से आगे, GPT-5.6 Sol के करीब), Terminal-Bench 4.0 20.3% → 38.0%। यह मुश्किल tasks पर ज़्यादा देर टिकता है और अपना काम ज़्यादा ध्यान से check करता है।
- Independent फैसला: बेहतर agent, वही मोहल्ला। Artificial Analysis इसे Intelligence Index पर 46 (4.6 से +2) और Coding Agent Index पर 56 (+9) देता है – कुल 4th स्थान, Fable 5.1, Astra और Opus 5 के पीछे।
- पेंच है token की भूख। Max reasoning effort पर यह प्रति task ~81K output tokens जलाता है – rivals से 2–3x – इसलिए $2/$6 की sticker price प्रति finished task लगभग $10/$50 flagships जितनी ही पड़ती है।
-
4.6 वाली ही कीमत: 200K tokens के नीचे $2 प्रति million input / $6 output
(उससे ऊपर दोगुना), cached input $0.50, और 2x कीमत पर 2x-speed fast variant। Cursor, Grok
Build (free try), Grok API (
grok-4-7), routers और clouds पर live।
Grok 4.7 क्या है
Grok 4.7 xAI का नया flagship है, Grok 4.6 (12 अगस्त को ship हुआ) का successor। यह closed, cloud-only model है – weights download नहीं होते – 500K-token context window, text और image input, text output, June 2026 pretraining cutoff और अगस्त 2026 तक का supplemental डेटा के साथ। API चार reasoning efforts देता है – low, medium, high (default) और xhigh – और कई headline benchmarks xhigh पर report किए गए हैं, इसलिए हर score को उसके effort label के साथ पढ़ें।
xAI के अनुसार इसमें तीन चीज़ें गई हैं: एक बिल्कुल नया बड़ा base model (2.1T figure), verifiable tasks के मुश्किल mix पर लंबा RL cycle जो कई घंटों वाले काम की ओर झुका है, और conversational और general knowledge work के लिए Grok Bot harness पर native training। SpaceX corpus अनोखा ingredient है – Musk का दावा है कि यह 4.7 को "real-world engineering" में सबसे अच्छा model बनाता है, जो डेटा की uniqueness का दावा है, कोई ऐसी चीज़ नहीं जिसे कोई public benchmark सीधे मापता हो।
रास्ता ऊबड़-खाबड़ था – और Musk ने pitch बदल दी
यह जानना ज़रूरी है, क्योंकि hype ने launch से ज़्यादा उम्मीदें लगा दी थीं। Musk ने जुलाई में 4.7 को "हर तरह से 4.6 से बेहतर, बस serve में थोड़ा slower, लेकिन token efficiency और भी बेहतर" बताया, फिर 12 अगस्त को कहा कि यह "सभी current models से आगे निकलेगा" और 3–4 हफ्तों में आएगा। 2 सितंबर को उन्होंने "10 दिन" कहा – यानी 12 सितंबर – और वह तारीख बिना release के निकल गई।
Delay की वजह जो Musk ने बताई, वह सच में दिलचस्प है: model मुश्किल tasks को बहुत जल्दी खत्म कर रहा था और अपना काम ठीक से check नहीं कर रहा था, क्योंकि RL में response length को बहुत ज़्यादा penalize किया गया था – reward ठीक करने के लिए "थोड़े और दिन पकने" की बात, कोई missing feature नहीं। और दावा धीमा हो गया: launch से दिनों पहले Musk ने 4.7 को "लगभग Opus 5.0 के बराबर, कुछ में बेहतर, कुछ में खराब" बताया, multimodal और image handling में अभी काम बाकी – और इससे आगे Grok 4.8 ("noticeable improvement," अक्टूबर की शुरुआत), 4.9 ("Astra या Fable class") और Grok 5 ("शायद सबसे बेहतर," जहाँ उन्हें AGI की उम्मीद है) की ओर इशारा किया। इनमें से किसी की कोई date या specs नहीं। मेरा निष्कर्ष: roadmap को signalling समझें, और 4.7 को नीचे दिए shipped numbers पर जज करें।
Benchmark नंबर जो मायने रखते हैं
पहले xAI की अपनी comparison table (Grok 4.7 vs 4.6, GPT-5.6 Sol, Fable 5.1), हर score के effort level के साथ – क्योंकि इस table में high vs xhigh असली फर्क डाल रहा है:
| Benchmark | क्या मापता है | Grok 4.7 | Grok 4.6 | Table में best rival |
|---|---|---|---|---|
| CursorBench 4.0 | लंबे coding tasks | 46.3% (xhigh) | 40.4% (high) | Astra-class (नीचे देखें) |
| DeepSWE v1.1 | Repo-scale software engineering | 71.0% (high) | 65.2% | 72.7% (Sol max) |
| Terminal-Bench 4.0 | कई घंटों का terminal work | 38.0% (xhigh, Grok Build) | 20.3% (high) | 57.9% (Astra) |
| SWE-Marathon v1.1 | Marathon coding sessions | 46.0% (high) | 31.9% | – |
| EEBench | Electrical-engineering tasks | 64.0% | – | – |
| AA Briefcase v1.1 | Professional knowledge work | 1,657 Elo | – | Opus 5 / Fable 5.1 |
| Harvey Legal Agent | Legal agent workflows | 19.6% | – | Frontier-class |
| HealthBench Professional | Clinical professional tasks | 56.7% (xhigh) | 48.5% | – |
DeepSWE वाली row ने मुझे चौंकाया: high effort पर 71.0% Fable 5.1 max (70.0%) को हराता है और Sol max (72.7%) के करीब है – और यह xhigh नंबर भी नहीं है। Terminal-Bench का लगभग दोगुना होना (20.3% → 38.0%) "ज़्यादा देर काम करता है" की कहानी एक figure में कहता है। लेकिन ध्यान दें कि xAI ने क्या headline नहीं बनाया: 38.0% के साथ Terminal-Bench 4.0 पर यह अब भी Astra के 57.9% और Fable 5.1 के 55.8% से काफी पीछे है, और Harvey legal score (19.6%) मामूली है। यह एक मज़बूत coding-and-documents release है, हर मोर्चे पर takeover नहीं।
अब independent check। Artificial Analysis ने Grok 4.7 (xhigh) को Grok 4.6 (high) के खिलाफ अपने Intelligence Index पर evaluate किया: 46 vs 44 – +2 की बढ़त जो SpaceXAI को top-4 labs में लाती है, लेकिन Fable 5.1 और Astra से पीछे रखती है। बढ़त ठीक वहीं है जहाँ xAI ने दावा किया था: AA-Briefcase 1,657 Elo (+111, Opus 5 और Fable 5.1 के ठीक पीछे, analytical quality 1,994 Elo के दम पर), GDPval-AA 1,695 (+90), Terminal-Bench 4.0 +4.5pp, GDP.pdf +3.0pp – साथ में AA-LCR (−3.7pp) और AutomationBench-AA (−1.1pp) पर गिरावट। Coding Agent Index (Grok Build harness के साथ) पर यह 47 → 56 छलांग लगाता है, सिर्फ Fable 5.1, Astra और Opus 5 के पीछे 4th स्थान पर, तीनों components में सुधार के साथ (DeepSWE 65% → 73%, Terminal-Bench 18% → 33%, SWE-Atlas-QnA 58% → 63%)। मोटे तौर पर vendor table और independent rerun सहमत हैं – जो ज़्यादातर launch weeks के बारे में नहीं कहा जा सकता।
Safety: अब तक का सबसे मज़बूत Grok, invite-only धार के साथ
xAI का कहना है कि 4.7 बिल्कुल नए safeguard stack के साथ बना है, और नंबर इतने specific हैं कि गंभीरता से लिए जा सकते हैं: यह LatchBio के biosafety benchmark में 62.4% के साथ top पर है (benign biology tasks पर utility + खतरनाक tasks पर refusal), और HackerBench v0.3 – risky और malicious cyber tasks – पर यह सिर्फ 3.3% risky dual-use prompts को जाने देता है, जबकि legitimate security work को शायद ही कभी block करता है। Hallucination में भी independent सुधार है: AA-Omniscience hallucination rate 4.6 के 34% से 29%, लगभग flat accuracy (47% vs 48%) पर।
देखने वाली बात: xAI चुनिंदा cybersecurity partners को defense research के लिए 4.7 की red-team capabilities का invite-only access देने लगा है। यह अब हर frontier lab का playbook है – safe model public, धारदार edges private। अगर आप internet से जुड़ा कुछ भी चलाते हैं, तो practical निष्कर्ष Astra वाला ही है: patch windows दोनों दिशाओं में छोटी हो रही हैं।
कीमत और उपलब्धता – ईमानदार गणित
Sticker price Grok 4.6 वाली ही है, और flagship के हिसाब से सच में सस्ती है:
- API: 200K prompt tokens के नीचे $2 प्रति million input / $6 output; उससे ऊपर पूरा request $4/$12 पर bill होता है। Cached input $0.50। Fast variant 2x कीमत पर 2x output speed देता है।
- Free try: Grok Build में free access; साथ ही Cursor, third-party coding harnesses, model routers और cloud platforms पर। API model ID
grok-4-7। - Context: 500K tokens, 4.6 जितना ही।
अब वह हिस्सा जो launch posts छोड़ देते हैं: throughput cost। Artificial Analysis ने Intelligence Index tasks पर 4.7 (xhigh) के ~81K output tokens मापे, बनाम 4.6 के ~36–38K और Astra max के ~27K – 125–196% ज़्यादा। इस volume पर प्रति task अनुमानित cost Grok 4.7 के लिए ~$3.74 बनाम Astra के ~$3.26 है, 5x sticker gap के बावजूद। Output speed ~188 tokens/second, प्रति task ~7.1 मिनट। तो "comparable models से दोगुना fast, आधी कीमत" प्रति token सच है और प्रति मिनट लगभग सच – लेकिन प्रति finished task token की भूख discount खा जाती है। अगर आप इस पर agents बनाते हैं, तो tokens पर नहीं tasks पर budget बनाएँ – हमारा AI API cost calculator आपके mix से ठीक यही गणित करता है।
अगर आप local AI चलाते हैं तो इसका क्या मतलब है
वह सवाल जो इस site के readers के मन में है: क्या Grok 4.7 मेरे RX 6800M के लिए कुछ बदलता है? तीन सटीक जवाब।
जहाँ gap बढ़ा: long-horizon agentic coding। DeepSWE 71%, SWE-Marathon 46%, CursorBench 46.3% – ये "Friday शुरू करो, Monday review करो" वाले workloads हैं जिन्हें 12 GB VRAM पर कोई 27B–35B model छू नहीं सकता। अगर आपका workflow multi-hour repo work delegate करना है, तो cloud और आगे निकल गया है, और Grok Build का free try आपके अपने repos पर verify करना सस्ता बनाता है।
जहाँ कुछ नहीं बदला: privacy, offline और छोटे scale पर प्रति-task cost। Local Qwen या Tiel-Coder MoE बिजली की cost पर चलता है और हर byte आपकी machine पर रखता है; Grok को आपका code किसी और के servers पर चाहिए। और छोटा, well-scoped task जो local 32B model एक pass में पूरा कर दे, अब भी billed API call के मुकाबले effectively free है। हमारे lab benchmarks दिखाते हैं कि local models असल में किसमें अच्छे हैं: fast, private, repeatable generation।
दिलचस्प बात: Grok 4.7 की delay story RL reward design का free lesson है – length को बहुत hard penalize करो तो model जल्दी quit करता है और अपना काम check करना बंद कर देता है। अगर आप अपने छोटे models को fine-tune या RL करते हैं, तो base model को दोष देने से पहले यह failure mode ("मुश्किल tasks बहुत जल्दी खत्म करना") check करने लायक है। Frontier labs की गलतियाँ open-source education हैं।
Bottom line
Grok 4.7 एक real, well-measured कदम है – लंबे coding और professional knowledge work पर अब तक का सबसे अच्छा Grok, सच में मज़बूत safety numbers के साथ – लेकिन यह वह "सभी current models से आगे" वाला release नहीं है जिसका preview Musk ने अगस्त में दिया था। यह लगभग Opus-5-class है, independent coding index पर 4th, terminal work पर Astra और Fable 5.1 से अब भी काफी पीछे, और इसके सस्ते tokens के साथ भारी token आदत आती है। इसे प्रति million tokens नहीं, प्रति finished task जज करें – और देखें कि क्या अक्टूबर में 4.8 वाकई gap बंद करता है।
Sources
- xAI / SpaceXAI: Introducing Grok 4.7 (21 सितंबर 2026)
- SpaceXAI Docs: Grok 4.7 technical overview – 500K context, pricing, reasoning effort
- Artificial Analysis: Benchmarking Grok 4.7 – Intelligence Index 46, Coding Agent Index 56
- heise online: Grok 4.7 – inexpensive API meets high token consumption
- TestingCatalog: SpaceXAI releases Grok 4.7 for coding and knowledge work