SpaceXAI shipped Grok 4.7 on 21 September 2026. The company calls it the most capable Grok model so far, and a clear step up from Grok 4.6 on long coding jobs, terminal work, and engineering tasks, at the same listed price and speed as 4.6. The model card is more useful than the launch post: it shows where the model leads, where it still sits behind Claude Fable 5.1 and GPT-6 Astra, and which products you can open today.
This is a launch note for people who build software. It sticks to the 21 September model card, the same-day product rollouts, and the public rate card. If a number comes from press coverage rather than the model card, it is marked as such.
What Grok 4.7 is
Grok 4.7 is a text model. It takes natural-language text and images as input and returns text. SpaceXAI is the public name of XAI LLC; the model card uses both names for the same company.
Training followed the usual stack: pretraining, then a longer supplemental stage than Grok 4.6, then supervised fine-tuning and reinforcement learning. The pretraining data cutoff is June 2026. Supplemental training uses data generated as late as August 2026. That supplemental pass includes anonymized Cursor workflow data, which is the part that matters if you use Grok inside an editor. Agentic reinforcement learning covered knowledge work, general coding, and purpose-built environments for kernel work, web development, and CAD.
Partner listings, including Vercel AI Gateway, put the context window at 500,000 tokens and describe four reasoning levels: low, medium, high, and xhigh. High is the default. A larger window helps when an agent has to hold a repo, a log, and a spec in one job. It does not remove the need to point the model at the right files.
Decrypt reported a parameter count of 2.1 trillion, about 40% above the 1.5 trillion it attributed to Grok 4.6. That figure is not in the model card text, so treat it as press reporting until SpaceXAI prints it on the card.
Where Grok 4.7 is live
The model card lists these channels as available with the 21 September release:
- SpaceXAI API at console.x.ai. The model id is
grok-4.7, on the standard chat and completions endpoints. - Grok Build, SpaceXAI’s terminal coding agent. Grok 4.7 is the default, through the API and the CLI.
- Cursor, for every user, on every plan.
- Office add-ins. Grok 4.7 is the default model in the Grok by SpaceXAI add-ins for Word, PowerPoint, and Excel.
- Model gateways, including OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic.

GitHub turned it on the same day. The GitHub Changelog for 21 September 2026 says Grok 4.7 is rolling out in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise, in Visual Studio Code, Visual Studio, the Copilot CLI, the cloud agent, the Copilot app, JetBrains, Xcode, and Eclipse. Rollout is gradual. Business and Enterprise admins control it from the Copilot model policy. Billing is provider list pricing under usage-based billing.
One point to check before you rely on it in a client conversation: the model card says consumer surfaces — the website, the mobile apps, and Grok inside X — are planned for a later date. Same-day reporting, including Decrypt, said the Grok app was already live with no waitlist. Those two statements were published on the same day. This draft leaves both on the record so the published version can match whatever is actually switched on.
What it costs
SpaceXAI’s launch post, quoted by Decrypt, called Grok 4.7 “a notable improvement over Grok 4.6 at the same price and speed.” Elon Musk described it on X as “a strong combination of intelligence, speed & low cost.”
The rate card published with the launch, in US dollars per million tokens:
| Prompt size | Input | Cached input | Output |
|---|---|---|---|
| Below 200,000 tokens | $2.00 | $0.50 | $6.00 |
| 200,000 tokens or more | $4.00 | $1.00 | $12.00 |
The long-context rates apply to the whole request once the prompt crosses 200,000 tokens. A job that looks cheap at 180,000 tokens can cost twice as much at 220,000. If you are moving a product from Grok 4.6, pin grok-4.7, run one representative task, and read the token bill before you switch the default.
Vercel is running a separate promotion: 40% off spacexai/grok-4.7 on AI Gateway, fx, and eve through 27 September 2026. That discount is a gateway offer with an end date, not a change to SpaceXAI’s list price.
What the model card actually scores
SpaceXAI’s own card is blunt about the shape of the gain. Grok 4.7 is built to finish longer tasks, check its own work, and spend fewer steps than earlier Grok models. On several coding charts it is still behind the current leaders. The graphs below use the card’s numbers. Each axis starts at zero.

Coding
- CursorBench 4.0 — long-horizon tasks from real Cursor sessions. Grok 4.7 scores 46.3% at xhigh and 43.9% at high. Version 4.0 is not comparable with CursorBench 3.2. The card does not print a full comparison table for this benchmark, so it is left as figures in the text.
- DeepSWE v1.1 — 113 original tasks across 91 repositories. Grok 4.7 scores 71.0% at high. On the same chart, GPT-5.6 Sol is at 72.7%, Fable 5 at 69.7%, Grok 4.6 at 65.2%, and Sonnet 5 at 54%.
- Terminal-Bench 4.0 — 66 hard terminal workflows, up to eight hours. Grok 4.7 scores 38.0% at xhigh in the Grok Build harness. Grok 4.6 is at 20.3% (high). Fable 5.1 leads that chart at 57.9%.
- FrontierSWE V2 — 34 ultra-long tasks, up to 20 hours, scored as partial credit. Grok 4.7 is at 29.0% (xhigh). Fable 5.1 is at 56.3%. Grok 4.6 is at 25.3%.
- SWE-Marathon v1.1 — 20 multi-hour engineering tasks, resolved only when every verifier passes. Grok 4.7 scores 46.0% at high. Opus 5 is at 50.0%. Grok 4.6 is at 31.9%.


The pattern is consistent. Against Grok 4.6, the jump is real, especially on terminal work (38.0% versus 20.3% on Terminal-Bench, at different effort levels) and on DeepSWE (71.0% versus 65.2% at the same effort). Against Fable 5.1 on the longest coding suites, the gap is still wide. FrontierSWE V2 is the clearest example: 29.0% for Grok 4.7 and 56.3% for Fable 5.1.
Knowledge work and engineering
- Legal Agent Benchmark — Vals AI’s held-out 120-task run, internet off. Grok 4.7 scores 19.6% at xhigh, ahead of Grok 4.6 at 15.8% and ahead of the other models on that chart. Absolute pass rates on this benchmark are low for everyone. The lead is real, and the ceiling is still low. Those two Grok scores use different effort levels, so they are kept in the text rather than forced onto the matched chart above.
- EEBench — circuits and chip-design tasks graded for physical correctness. Grok 4.7 scores 66.0% at xhigh, against 60.0% for Grok 4.6 at the same effort. GPT-6 Astra leads at 69.3%. Opus 5 is at 61.6%.
- CADGenBench — CAD geometry from a design prompt. Grok 4.7 scores 44.4% at high, against 40.9% for Grok 4.6 at the same effort.

Decrypt’s write-up of the wider leaderboards matches that picture. It reported a GDPval Elo of 1695 for Grok 4.7 against 1735 for Claude Fable 5.1, and an AA-Briefcase Elo of 1657 against 1678 for Fable 5.1. Days before launch, Musk wrote that 4.7 should land roughly level with Claude Opus 5.0, and that multimodal performance still needed work. He also sketched Grok 4.8, Grok 4.9, and Grok 5, with no dates. Those Elo figures are press reporting of leaderboards, so they are not drawn on the charts above.
What this changes if you ship software

For day-to-day coding in Cursor, the practical change is access. Grok 4.7 is on every Cursor plan, and the model was trained with anonymized Cursor workflow data specifically to improve coding and agent behaviour. That is a reason to run it on a real ticket — a bug with a failing test, a migration, a messy plugin — and compare the diff with the model you already trust. A demo prompt will not show you the Terminal-Bench gap or the DeepSWE gain.
For an API product, the sticker price matches Grok 4.6 on short prompts. The decision is whether the higher scores on long agent jobs are worth the 200,000-token price step, and whether your users are better served by this model or by whichever frontier model already wins your own eval. The model card is explicit about one limit: Grok 4.7 is for engineering and creative work with a person in the loop when the decision is medical, legal, financial, or safety-critical.
For office work, the Word, PowerPoint, and Excel add-ins now default to 4.7. That lines up with the knowledge-work training, and with a legal-agent score that moved up while staying modest in absolute terms. Use it to draft and structure. Keep a human on anything a client will sign.
A sensible way to try it this week
- In Cursor, pick Grok 4.7 on one real branch. Give it the failing test and the files that own the bug. Read the patch. That single run tells you more than the launch thread.
- If you call the API, send one production-shaped request to
grok-4.7and record input tokens, cached tokens, and output tokens. Confirm you are under or over the 200,000-token line before you estimate a monthly bill. - If Copilot is where your team already works, check the model picker after the gradual rollout, and check the org policy if you are on Business or Enterprise.
- Read the chart that matches the job. DeepSWE, Terminal-Bench, and EEBench do not crown the same model.
Frequently asked questions
-
When did Grok 4.7 launch?
SpaceXAI released Grok 4.7 on 21 September 2026. The model card is dated the same day. SpaceXAI is the public name of XAI LLC, and the card uses both names for the same company.
-
Where can you use Grok 4.7?
The model card lists the SpaceXAI API under the model id grok-4.7, Grok Build as the default model, Cursor for every user on every plan, and the Grok add-ins for Word, PowerPoint, and Excel as the default. It also lists gateways including OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic. GitHub began a gradual Copilot rollout the same day for Pro, Pro+, Max, Business, and Enterprise. The card says the Grok website, the mobile apps, and Grok inside X are planned for a later date. Same-day reporting, including Decrypt, said the Grok app was already live.
-
What does the Grok 4.7 API cost?
For prompts under 200,000 tokens, the launch rate card lists 2 US dollars per million input tokens, 0.50 per million cached input tokens, and 6 per million output tokens. Once a prompt reaches 200,000 tokens, those rates become 4, 1, and 12 dollars for the whole request. SpaceXAI described Grok 4.7 as the same price and speed as Grok 4.6. Vercel is discounting spacexai/grok-4.7 by 40 percent on AI Gateway through 27 September 2026. That discount is a gateway offer, separate from the list price.
-
How large is the Grok 4.7 context window?
Partner listings, including Vercel AI Gateway, put the context window at 500,000 tokens. The same listings describe four reasoning levels: low, medium, high, and xhigh, with high as the default. The model card sets the pretraining data cutoff at June 2026 and says supplemental training uses data generated as late as August 2026. The model accepts text and images and returns text.
-
Is Grok 4.7 stronger than Grok 4.6?
On the five model-card comparisons that use the same reasoning effort, Grok 4.7 scores higher on every one. SWE-Marathon v1.1 moves from 31.9 percent to 46.0 percent at high. DeepSWE v1.1 moves from 65.2 percent to 71.0 percent at high. EEBench moves from 60.0 percent to 66.0 percent at xhigh. CADGenBench moves from 40.9 percent to 44.4 percent at high. FrontierSWE V2 moves from 25.3 percent to 29.0 percent at xhigh.
-
Does Grok 4.7 lead current coding benchmarks?
It leads some charts in the model card and trails the leaders on others. On DeepSWE v1.1 it scores 71.0 percent, second to GPT-5.6 Sol at 72.7 percent. On Terminal-Bench 4.0, Fable 5.1 leads at 57.9 percent and Grok 4.7 scores 38.0 percent at xhigh. On EEBench, GPT-6 Astra leads at 69.3 percent and Grok 4.7 scores 66.0 percent. Decrypt reported a GDPval Elo of 1695 for Grok 4.7, against 1735 for Claude Fable 5.1.
-
Can developers use Grok 4.7 in Cursor and GitHub Copilot?
Yes. The model card says Grok 4.7 is available in Cursor for every user on every plan, and that the model received supplemental training on anonymized Cursor workflow data. GitHub’s changelog for 21 September 2026 says it is rolling out in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise, in Visual Studio Code, Visual Studio, JetBrains, Xcode, Eclipse, the Copilot CLI, the cloud agent, and the Copilot app. Business and Enterprise admins turn it on or off from the Copilot model policy. Rollout is gradual.
-
Can Grok 4.7 make medical, legal, or financial decisions on its own?
The model card says Grok 4.7 is not intended for autonomous high-stakes decisions in medicine, law, finance, or safety-critical systems unless a person with domain expertise checks the result. On the Legal Agent Benchmark, Grok 4.7 scores 19.6 percent at xhigh, which leads that held-out chart. The absolute pass rate on that benchmark is still low.
Sources
- SpaceXAI, Model Card: Grok 4.7, 21 September 2026. Scores, training cutoff, and availability channels in this article come from that card.
- GitHub Changelog, Grok 4.7 is now available in GitHub Copilot, 21 September 2026.
- Vercel Changelog, Grok 4.7 now available and 40% off on AI Gateway, fx, and eve, 21 September 2026.
- Decrypt, xAI Launches Grok 4.7, 21 September 2026. Used here for the launch-post wording, the parameter-count report, the GDPval and AA-Briefcase figures, and Musk’s pre-launch note on Opus 5.0.
Bdeb Technology builds websites, Android apps, and custom systems. If you want a working product on top of a model like this — a chatbot, an internal tool, or a workflow that has to survive contact with real users — that is the work. A written quote comes back within 24 hours.


Leave a Reply