SpaceXAI (formerly xAI) released Grok 4.6 on August 12, 2026 — just one month after Grok 4.5. The new model matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index and focuses on long-running agents and complex multi-step tasks. It’s immediately available in Cursor and Grok Build, with doubled usage limits for the first week.

What Happened

Grok 4.6 achieves a score of 61 on the AA Intelligence Index, tied with OpenAI’s GPT-5.6 Sol Max and five points ahead of Claude Fable 5 Max (56). The benchmark composites nine agentic coding and knowledge work tests including DeepSWE 1.1, CursorBench 3.2, and FrontierCode 1.1.

The model underwent a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and advanced technical concepts, plus high-quality engineering data. SpaceXAI also improved the optimizer and training recipe. The training pipeline included supervised fine-tuning (SFT) using Grok 4.5’s outputs, followed by reinforcement learning.

SpaceXAI is offering 2x included usage inside Grok Build and Cursor for the first week to encourage testing. The release comes after SpaceXAI’s parent company SpaceX completed its acquisition of xAI and went public via the largest IPO on record in June.

Read the full announcement →

My Take

This is a surprisingly aggressive release cadence from SpaceXAI. Shipping a major model update one month after 4.5 signals they’re in an all-out sprint to match — or exceed — OpenAI’s pace. The focus on long-running agents and multi-step coding tasks directly targets the developer market, where Cursor integration gives them a ready-made distribution channel.

The AA Intelligence Index tie with GPT-5.6 is notable, but benchmarks only tell part of the story. The real test is whether Grok 4.6 can maintain coherence and correctness across hours-long agent sessions. If the supplemental training on “high-quality engineering data” translates to fewer hallucinations in complex codebases, this could be the model that makes agentic coding workflows practical for production use.

What to Watch

  • Whether Grok 4.6’s performance holds up in real-world Cursor usage beyond benchmarks — especially on long-running refactoring tasks.
  • How quickly SpaceXAI expands availability beyond Cursor and Grok Build to API access and other IDEs.
  • The competitive response from OpenAI and Anthropic — we may see GPT-5.7 or Claude Fable 5.1 within weeks.