Grok 4.7 Launches With $2/$6 API Pricing, 500K Context and Major Coding Gains

Grok 4.7 Launches With $2/$6 API Pricing, 500K Context and Major Coding Gains

SpaceXAI has launched Grok 4.7, a new frontier model aimed at coding, long-horizon agentic work and professional knowledge tasks, while keeping the same base API pricing as Grok 4.6.

The September 21 release is positioned as a substantial upgrade rather than a minor refresh. SpaceXAI says Grok 4.7 uses a new, larger base model, was trained with a longer reinforcement-learning run on a harder task mix, and was optimized for work that can take hours rather than seconds. The company also says the model is better at checking its own work, managing long context and operating inside agentic coding environments.

For developers, the headline economics remain aggressive: $2 per million input tokens and $6 per million output tokens for the standard Grok 4.7 API model. That is the same starting token price SpaceXAI charged for Grok 4.6, despite the company reporting higher scores across every benchmark in its launch comparison against the previous generation.

Grok 4.7 is available now through the xAI API, Grok Build, Cursor and several model gateways. SpaceXAI's developer documentation identifies the public API model as grok-4.7 and lists a 500,000-token context window, text and image input, text output and four selectable reasoning levels: low, medium, high and xhigh.

What is Grok 4.7?

Grok 4.7 is SpaceXAI's latest general-purpose frontier model for software engineering, agentic workflows and professional knowledge work. It follows Grok 4.6, which launched in August 2026, and continues the company's recent focus on models that can operate for longer periods inside coding and tool-use environments.

SpaceXAI describes the model as more capable at difficult tasks that require sustained work, self-verification and context management. According to the company, training placed more weight on problems that take many hours to complete rather than short benchmark-style prompts. That shift matters because many of the most commercially valuable AI tasks increasingly involve long sequences of actions: reading a repository, planning a change, editing multiple files, running tests, debugging failures, reviewing output and iterating until the job is complete.

The company also says Grok 4.7 was trained to natively understand the Grok Bot harness. In practical terms, that means the model was tuned not only for standalone question answering but also for the surrounding agent system in which it plans, calls tools, reacts to environment feedback and continues work across multiple steps.

That helps explain why the launch materials emphasize coding agents, office work and terminal tasks more than traditional academic benchmarks.

Grok 4.7 pricing stays at $2 input and $6 output per million tokens

SpaceXAI is keeping the standard Grok 4.7 API price at the same level as Grok 4.6:

  • Input: $2 per 1 million tokens
  • Output: $6 per 1 million tokens

The launch comparison places those rates beside $4 input and $20 output for GPT-5.6 Sol Max and $10 input and $50 output for Fable 5.1 Max. Those figures are part of SpaceXAI's own launch table and should be interpreted alongside differences in reasoning effort, serving configuration and benchmark methodology rather than as a universal apples-to-apples cost ranking.

SpaceXAI also offers a faster serving option. Its documentation says Grok 4.7 Fast uses the same model on faster infrastructure and costs twice the standard token rate. The Fast variant is available in Cursor and Grok Build but is not currently exposed through the public xAI API.

For API developers, this pricing puts Grok 4.7 in a position where cost can become as important as raw benchmark performance. A model that is slightly behind a more expensive rival on one benchmark but materially cheaper in production may still be attractive for high-volume agent workloads, code-generation systems and automated back-office tasks.

Grok 4.7 benchmark results

SpaceXAI published a direct launch comparison between Grok 4.7 xhigh, Grok 4.6 high, GPT-5.6 Sol max and Fable 5.1 max. The company reports gains for Grok 4.7 over Grok 4.6 across every row in the table.

Benchmark Grok 4.7 Grok 4.6 GPT-5.6 Sol Fable 5.1
CursorBench 4.0 46.3% 40.4% 41.7% 51.8%
DeepSWE v1.1 71.0%* 65.2% 72.7% 70.0%
EEBench 64.0% 53.0% 39.4% 56.4%
AA Briefcase v1.1 1,657 1,546 1,487 1,678
Terminal-Bench 4.0 38.0% 20.3% 37.3% 57.9%
Harvey Legal Agent Benchmark 19.6% 15.8% 2.5% 6.7%
HealthBench Professional 56.7% 48.5% 60.5% 62.1%

*SpaceXAI marks the Grok 4.7 DeepSWE result as a high-effort run rather than xhigh.

The most important comparison is the one against Grok 4.6 because it measures the new model against the previous SpaceXAI generation under a closely related product environment. On that basis, the reported improvements are broad.

CursorBench rises from 40.4% to 46.3%. Terminal-Bench 4.0 climbs from 20.3% to 38.0%, which is one of the largest jumps in the table. EEBench moves from 53.0% to 64.0%, while Harvey's Legal Agent Benchmark rises from 15.8% to 19.6%. HealthBench Professional improves from 48.5% to 56.7%.

Those results suggest the upgrade is not confined to coding. SpaceXAI is specifically arguing that Grok 4.7 is better at professional work spanning engineering, legal tasks, office workflows and clinical reasoning.

How Grok 4.7 compares with GPT-5.6 Sol and Fable 5.1

The launch table shows a mixed picture rather than a clean sweep.

Against GPT-5.6 Sol Max, Grok 4.7 posts a higher score in SpaceXAI's table on CursorBench 4.0, EEBench, AA Briefcase, Terminal-Bench 4.0 and the Harvey Legal Agent Benchmark. GPT-5.6 Sol leads on DeepSWE v1.1 and HealthBench Professional.

Against Fable 5.1 Max, Grok 4.7 leads on DeepSWE, EEBench and Harvey's legal benchmark, while Fable leads on CursorBench, AA Briefcase, Terminal-Bench and HealthBench Professional.

That is why the release is better described as a major price-performance push than as proof that one model is universally superior. Different benchmark families measure different capabilities, and the models are shown at different effort settings. Independent evaluations may also produce different rankings once Grok 4.7 has been tested outside SpaceXAI's launch environment.

Still, the pricing gap is notable. SpaceXAI is presenting Grok 4.7 at $2 input and $6 output per million tokens while showing it competitive with models it lists at materially higher token prices. For teams deploying large agent fleets, that can translate into meaningful operating-cost differences.

CursorBench and long-running software engineering

SpaceXAI highlights CursorBench 4.0 prominently because the benchmark is designed around longer-running software-engineering work rather than isolated code snippets.

Grok 4.7 scores 46.3% in the company's launch table, up from 40.4% for Grok 4.6. Fable 5.1 remains ahead at 51.8%, while GPT-5.6 Sol is listed at 41.7%.

The result supports SpaceXAI's broader training claim: Grok 4.7 was optimized for tasks that require sustained execution. In real software development, the hard part is often not writing a single function. It is understanding an unfamiliar codebase, identifying dependencies, planning changes, editing safely, running tests and continuing after failures.

That is the kind of workload frontier coding agents are increasingly expected to handle.

Terminal-Bench nearly doubles over Grok 4.6

The biggest generation-over-generation jump in the headline comparison comes on Terminal-Bench 4.0.

Grok 4.6 is listed at 20.3%. Grok 4.7 reaches 38.0%.

Terminal benchmarks test a model's ability to operate inside command-line environments where success depends on sequencing actions correctly, interpreting tool output and recovering from errors. Those skills are central to autonomous coding agents and infrastructure assistants.

Grok 4.7 narrowly edges the GPT-5.6 Sol figure shown in the table at 37.3%, though Fable 5.1 remains far ahead at 57.9%.

The jump from the previous Grok generation is more important than the one-point difference against GPT-5.6 Sol. It indicates that SpaceXAI's training changes may be translating into stronger tool-use behavior rather than only better static reasoning.

EEBench is one of Grok 4.7's strongest rows

On EEBench, which measures electrical-engineering capability, Grok 4.7 is listed at 64.0%.

That compares with 53.0% for Grok 4.6, 39.4% for GPT-5.6 Sol and 56.4% for Fable 5.1 in SpaceXAI's table.

The result fits the company's effort to broaden Grok beyond coding into engineering and scientific work. A frontier model that can reason about hardware, circuits and engineering design has potential use cases in product development, diagnostics, documentation, verification and technical research.

As always, a benchmark score does not establish reliability for safety-critical engineering on its own. Human review, simulation, testing and domain-specific validation remain necessary in real deployments.

Legal and professional knowledge work

SpaceXAI also emphasizes Grok 4.7's performance on professional tasks that look more like real workplace projects than traditional question-answer tests.

On the Harvey Legal Agent Benchmark, the model scores 19.6% in the company's launch comparison, up from 15.8% for Grok 4.6. The same table lists GPT-5.6 Sol at 2.5% and Fable 5.1 at 6.7%.

The absolute numbers are a useful reminder that complex end-to-end legal work remains difficult for frontier AI. Harvey's own benchmark research has noted that strict all-pass legal-agent tasks are far from saturated. A higher result is meaningful, but it does not imply that a model can replace legal professionals or operate without review.

AA Briefcase is another important signal because it tests multi-hour office work across professional domains. Grok 4.7 reaches 1,657 in SpaceXAI's table, ahead of Grok 4.6 at 1,546 and GPT-5.6 Sol at 1,487, while Fable 5.1 is slightly higher at 1,678.

SpaceXAI says Grok 4.7 is also better at creating documents and presentations, tying the model more directly to knowledge-worker use cases rather than treating coding as the only target market.

Clinical reasoning improves, but rivals remain ahead

Grok 4.7 scores 56.7% on HealthBench Professional in the launch materials, compared with 48.5% for Grok 4.6.

GPT-5.6 Sol is listed at 60.5%, and Fable 5.1 at 62.1%.

This is another area where the release shows clear internal improvement without leading the comparison. It also reinforces the need for caution when interpreting healthcare benchmarks. Clinical reasoning evaluations can measure useful capabilities, but they are not a substitute for medical validation, professional oversight or regulatory requirements.

500,000-token context window

SpaceXAI's developer documentation lists Grok 4.7 with a 500,000-token context window.

That is large enough for many repository-scale coding tasks, extensive document collections and long agent sessions. The model supports text and image input with text output and, according to the documentation, has no fixed text-output limit listed on the model page.

The knowledge cutoff is listed as May 2026. For current information, developers can combine the model with SpaceXAI's web search and X search tools rather than relying only on its internal training data.

The model also supports function calling and code execution, making it suitable for tool-using agents rather than simple chat-only deployments.

Reasoning effort can be set from low to xhigh

Developers can control Grok 4.7's reasoning effort using four settings:

  • low
  • medium
  • high
  • xhigh

High is the documented default. The xhigh setting allows the model to spend more effort on difficult problems, but developers should expect a tradeoff between latency, token consumption and accuracy.

This also matters when reading benchmark tables. Grok 4.7 is shown at xhigh on most of SpaceXAI's comparison rows, while the DeepSWE score is specifically marked as high effort. Rival models are shown at their own max settings.

Reasoning configuration can materially change both benchmark performance and production economics, so teams should test models using the settings they actually plan to deploy.

New base model and longer reinforcement-learning run

The architectural story behind Grok 4.7 is one of scale and post-training rather than a new user-facing modality.

SpaceXAI says the model uses a larger base model than Grok 4.6 and received a longer reinforcement-learning run on a harder task distribution. That task mix was weighted toward work requiring many hours of sustained reasoning and tool use.

The company also emphasizes self-verification. A model that can detect its own mistakes before committing to an answer or code change is potentially more valuable than one that simply generates stronger first-pass output.

For coding agents, self-checking can mean reading test failures and revising a patch. For office work, it can mean comparing a completed document against instructions before finishing. For engineering tasks, it can mean verifying calculations or checking whether assumptions remain consistent across a long workflow.

Those behaviors are difficult to capture with short-form benchmarks, which is why SpaceXAI is putting more emphasis on long-horizon evaluations.

Grok 4.7 is now the default model in Grok Build

SpaceXAI says Grok 4.7 is now the default model in Grok Build, the company's coding-agent environment.

It is also available in Cursor on all plans, according to the developer documentation. Cursor separately announced the model's availability and highlighted the same reported gains on CursorBench and Terminal-Bench.

For developers who do not want to integrate directly with the xAI API, the model is also available through model gateways including OpenRouter, Vercel and Cloudflare.

The public API model identifier is:

grok-4.7

SpaceXAI also offers a United States regional endpoint for teams that want inference to remain in the U.S. The documentation says token usage through that endpoint carries a 10% premium.

Encrypted reasoning in the Responses API

One technical change highlighted in the Grok 4.7 documentation is how multi-turn reasoning is handled in the Responses API.

Responses from grok-4.7 include encrypted reasoning content that can be passed back unchanged in subsequent requests. The goal is to preserve reasoning continuity across multi-turn conversations without exposing the underlying private chain of thought.

For agent developers, that can make long workflows easier to maintain because the model can continue from its prior reasoning state while the application only handles encrypted reasoning objects.

SpaceXAI also recommends using a prompt cache key so related conversation requests are routed in a way that improves cache-hit reliability. For long agent loops, the documentation recommends context compaction as well.

Grok 4.7 safety and cybersecurity changes

SpaceXAI says Grok 4.7 ships with a newly designed safeguard stack and is its strongest model so far on refusal behavior and jailbreak resistance.

The company reports a score of 62.4% on LatchBio's biosafety benchmark. It also says Grok 4.7 allowed only 3.3% of risky dual-use prompts through on its HackerBench v0.3 evaluation while maintaining a low refusal rate on legitimate security work.

These are company-reported launch figures and should be evaluated alongside independent safety research as it becomes available.

SpaceXAI says it is also giving selected cybersecurity partners invite-only access to Grok 4.7 red-team capabilities for defensive research.

The broader theme is calibration: the company is trying to make the model more capable on legitimate cyber and biological work while reducing compliance with requests that could enable serious misuse.

What changed from Grok 4.6?

From a buyer's perspective, the upgrade can be summarized in five areas.

  1. Larger model: SpaceXAI says Grok 4.7 uses a new and larger base model.
  2. Longer RL training: post-training focused more heavily on difficult, long-duration tasks.
  3. Better agent behavior: the model was trained around the Grok Bot harness and long tool-use workflows.
  4. Broad benchmark gains: every Grok 4.7 score in the launch table improves over Grok 4.6.
  5. No base price increase: standard API pricing remains $2 input and $6 output per million tokens.

That last point may be the most important commercially. Frontier model launches often arrive with higher capability and higher cost. SpaceXAI is instead trying to make Grok 4.7 a drop-in upgrade for teams already comfortable with Grok 4.6 economics.

What developers should test before switching

Benchmark improvements do not automatically mean every production application should migrate immediately.

Teams using Grok 4.6 should test Grok 4.7 against their own workloads, especially if they rely on structured outputs, long context, tool execution or deterministic coding behavior.

Useful migration tests include:

  • repository-level bug fixing and refactoring
  • long-running terminal workflows
  • function-calling reliability
  • tool selection accuracy
  • JSON or schema-constrained output
  • latency at high and xhigh reasoning
  • prompt-cache behavior
  • cost per successfully completed task
  • regression rates on established production prompts

The last metric is particularly important. The cheapest model by token is not necessarily the cheapest model per completed workflow. A more capable model can sometimes cost less overall if it requires fewer retries, shorter conversations and less human correction.

Why Grok 4.7 matters

The launch reflects a broader change in the frontier AI market.

Competition is moving away from one-shot chatbot intelligence and toward systems that can execute useful work over long periods. That means model vendors increasingly need to compete on more than reasoning benchmarks. They need to offer strong tool use, long context, reliable self-correction, predictable pricing and deployment infrastructure.

Grok 4.7 is clearly designed for that environment.

SpaceXAI is not positioning it only as a smarter chat model. The launch language repeatedly centers software engineering, professional knowledge work, terminal use, office tasks and agentic execution.

The 500,000-token context window, integrated tools, selectable reasoning levels and emphasis on long-duration RL training all support that strategy.

At the same time, rivals remain stronger on several benchmark rows, and independent testing will be important before drawing broad conclusions about where Grok 4.7 sits across the frontier-model market.

The bottom line

Grok 4.7 is now live, and SpaceXAI is making a clear bet on price-performance for agentic work.

The model keeps standard API pricing at $2 per million input tokens and $6 per million output tokens, offers a 500,000-token context window and is available through the xAI API, Grok Build, Cursor and major model gateways.

SpaceXAI reports substantial improvements over Grok 4.6 across coding, terminal work, engineering, legal tasks, office work and clinical reasoning. The company attributes those gains to a larger base model, a longer reinforcement-learning run and training focused more heavily on difficult tasks that require sustained execution.

The launch comparison also shows Grok 4.7 trading wins with GPT-5.6 Sol and Fable 5.1 rather than dominating every category. That makes the pricing story particularly important: SpaceXAI is offering competitive frontier performance at a lower listed token cost than the two rival configurations shown in its launch materials.

For developers, the most important question will not be whether Grok 4.7 wins a single benchmark. It will be whether the model completes real coding and knowledge-work tasks more reliably per dollar than the alternatives.

That is the test that begins now.


Sources:

SpaceXAI — Introducing Grok 4.7
SpaceXAI developer documentation — Grok 4.7
Cursor — Grok 4.7 is now live
Harvey — Legal Agent Benchmark methodology and initial results

'; slot.appendChild(frame); })();
Sponsored: View offer
Back to blog