xAI and Cursor Ship Grok 4.6: It Ties GPT-5.6 Sol on the Composite Score, but Still Trails by 7–8 Points on Two Hard Agent Benchmarks

On August 12, xAI and Cursor jointly released Grok 4.6, aimed at long-running agents and more ambitious interactive and visual work. In the vendors' own numbers it scores 61 on the Artificial Analysis Intelligence Index—a composite of nine benchmarks—matching GPT-5.6 Sol Max and sitting just below Fable 5 Max's 62. Broken out benchmark by benchmark, though, DeepSWE and Terminal-Bench remain clear weak spots. Pricing holds at $2/$6 per million tokens, and it went live the same day in Cursor, Grok Build, the xAI API, and via OpenRouter, Vercel, and Cloudflare.

Past the Composite, the Spread Is Wide

The wins in the official table cluster in knowledge work and front-end tasks. On GDPVal-AA v2 Grok 4.6 posts 1753, ahead of GPT-5.6 Sol Max's 1728 and Fable 5 Max's 1741. AA-Briefcase at 1577 is the highest of the three. On the legal benchmark Harvey LAB it scores 15.8%, well clear of the other two at 2.5% and 11.3%. The losses are just as clear: DeepSWE v1.1 comes in at 65.9% against GPT-5.6 Sol's 73% and Fable 5's 70%, and Terminal-Bench v3.0 lands at 26% against 34.6% and 34.1%. The generation-over-generation gains are real, though—Terminal-Bench went from 15.7% to 26%, and DeepSWE from 54% to 65.9%.

The Gains Come From Post-Training, Not Scale

By both companies' account, Grok 4.6 went through a longer supplemental training run than 4.5: curated model-generated reasoning data, high-quality engineering data, and an improved optimizer and recipe, producing a stronger foundation before the SFT and RL stages. For SFT, they used Grok 4.5 to regenerate trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, filtering out problematic traces with model-based checks. RL then spread across concrete environments—kernel optimization, web development, computer-aided design, and more. They also report that on longer trajectories the model began self-testing, verifying its own work before moving on.

How to Read These Numbers

All of them are vendor-run or self-reported, and no third-party reproduction on a neutral harness has landed yet. CursorBench carries an extra structural wrinkle: Cursor both maintains the benchmark and co-trains the model with xAI. At the previous release, Cursor itself disclosed that "an earlier snapshot of the Cursor codebase was accidentally included in training," giving Grok 4.5 an advantage on CursorBench, with "the exact impact unclear"—adding that the data was removed for future models and that the benchmark got a substantial overhaul, now at version 3.2. On CursorBench v3.2, Grok 4.6 scores 69.9%, still below Fable 5 Max's 70.5%.

Why It Matters

The positioning is legible. Pricing did not move ($2/$6 per million tokens, with a fast variant at double), the composite score puts it in the top tier, and long-running agents, interactive and visual output, and knowledge work are where it is strongest—with double usage inside Cursor and Grok Build for the first week if you want to try it. But if your actual work is repository-scale code repair and terminal-environment operations, that 7–8 point gap on DeepSWE and Terminal-Bench is still there. Run your own real tasks before switching, rather than deciding off the composite alone.

via: xAI's announcement, Introducing Grok 4.6; Cursor's companion post; Cursor on Grok 4.5 and CursorBench; verified 2026-08-13