
Introducing Grok 4.7
Our most capable model for long-running coding and knowledge work.
A frontier model for long-running coding and knowledge work, at the same price and speed as Grok 4.6.

Our most capable model for long-running coding and knowledge work.

The official model card for Grok 4.7. It documents the model's capability benchmarks and safeguard evaluations.

Cursor is partnering with SpaceX to accelerate our model training efforts.
Grok 4.7 stays with difficult tasks across many steps. It uses tools, checks its own work, adjusts its approach, and keeps moving toward a finished result.
Grok 4.7 works across large codebases and extended engineering projects. It researches unfamiliar systems, edits across files, runs tests, and verifies the result.
Beyond code, Grok 4.7 works across documents, spreadsheets, presentations, and other professional artifacts. Give it a multi-step project that requires research, analysis, and synthesis.
Grok 4.7 verifies results more carefully than Grok 4.6 and manages longer context, with a 256k standard window and 500k long context.
Grok 4.7 is built for long-running coding and knowledge work. It uses a larger base than Grok 4.6 and a longer training run on harder, multi-hour tasks. It handles projects that span research, analysis, implementation, and several rounds of refinement.
On CursorBench 4.0 it scores 46.3% at extra high effort. It is served at the same price and speed as Grok 4.6, and Cursor subscription plans for individuals and teams include significant usage of the model.
We trained Grok 4.7 jointly with SpaceXAI on a new, larger base model than Grok 4.6. The reinforcement learning run was longer and weighted toward problems that take many hours to complete.
The model is better at verifying its own work and managing longer context. It was also trained to understand the Grok Bot harness, which improves conversational tasks and general knowledge work.
| Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max | |
|---|---|---|---|---|
| CursorBench 4.0Software engineering | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1Software engineering | 71.0% high effort | 65.2% | 72.7% | 70.0% |
| EEBenchElectrical engineering | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1Multi-hour office work | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0Multi-hour terminal work | 37.6% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent BenchmarkLegal work | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench ProfessionalClinical reasoning | 56.7% | 48.5% | 60.5% | 62.1% |
Scores from the Grok 4.7 announcement. The DeepSWE result for Grok 4.7 is at high effort.
On GDPval, Grok 4.7 scores 1,695 Elo at extra high effort, compared with 1,605 for Grok 4.6, 1,735 for Fable 5.1, and 1,542 for GPT-6 Astra.
(Above) Grok 4.7 results across agentic coding and knowledge work benchmarks.
Grok 4.7 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

A frontier model for long-running coding and knowledge work, at the same price and speed as Grok 4.6.

Our most capable model for long-running coding and knowledge work.

The official model card for Grok 4.7. It documents the model's capability benchmarks and safeguard evaluations.

Cursor is partnering with SpaceX to accelerate our model training efforts.
Grok 4.7 stays with difficult tasks across many steps. It uses tools, checks its own work, adjusts its approach, and keeps moving toward a finished result.
Grok 4.7 works across large codebases and extended engineering projects. It researches unfamiliar systems, edits across files, runs tests, and verifies the result.
Beyond code, Grok 4.7 works across documents, spreadsheets, presentations, and other professional artifacts. Give it a multi-step project that requires research, analysis, and synthesis.
Grok 4.7 verifies results more carefully than Grok 4.6 and manages longer context, with a 256k standard window and 500k long context.
Grok 4.7 is built for long-running coding and knowledge work. It uses a larger base than Grok 4.6 and a longer training run on harder, multi-hour tasks. It handles projects that span research, analysis, implementation, and several rounds of refinement.
On CursorBench 4.0 it scores 46.3% at extra high effort. It is served at the same price and speed as Grok 4.6, and Cursor subscription plans for individuals and teams include significant usage of the model.
We trained Grok 4.7 jointly with SpaceXAI on a new, larger base model than Grok 4.6. The reinforcement learning run was longer and weighted toward problems that take many hours to complete.
The model is better at verifying its own work and managing longer context. It was also trained to understand the Grok Bot harness, which improves conversational tasks and general knowledge work.
| Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max | |
|---|---|---|---|---|
| CursorBench 4.0Software engineering | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1Software engineering | 71.0% high effort | 65.2% | 72.7% | 70.0% |
| EEBenchElectrical engineering | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1Multi-hour office work | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0Multi-hour terminal work | 37.6% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent BenchmarkLegal work | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench ProfessionalClinical reasoning | 56.7% | 48.5% | 60.5% | 62.1% |
Scores from the Grok 4.7 announcement. The DeepSWE result for Grok 4.7 is at high effort.
On GDPval, Grok 4.7 scores 1,695 Elo at extra high effort, compared with 1,605 for Grok 4.6, 1,735 for Fable 5.1, and 1,542 for GPT-6 Astra.
(Above) Grok 4.7 results across agentic coding and knowledge work benchmarks.
Grok 4.7 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

(Above) Grok 4.7 scores 46.3% at extra high effort on CursorBench 4.0.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.7 Extra High | 46.3% | $6.01 | 70,141 | 88 |
(Above) Grok 4.7 scores 46.3% at extra high effort on CursorBench 4.0.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.7 Extra High | 46.3% | $6.01 | 70,141 | 88 |