xAI launched Grok 4.7 on Monday, achieving the highest score on the Harvey Legal Agent Benchmark with 19.6%, compared to Fable 5.1's 6.7% and GPT-5.6 Sol's 2.5%.
The model is available on the Grok app, Cursor, Grok Build, and the xAI API without a waitlist.
Pricing and performance
Grok 4.7 is priced at $2 per million input tokens and $6 per million output tokens, matching Grok 4.6 rates. Input pricing is half of GPT-5.6 Sol's $4 and one-fifth of Fable 5.1's $10. Output costs are $6 compared to $20 for Sol and $50 for Fable.
On the CursorBench 4.0 cost-performance metric, Grok 4.7 achieved approximately 46% accuracy at roughly $6 per task. Fable 5.1 scored higher at 51.8% but required nearly $17 per task, while Claude Opus 5 needed roughly double the spending for similar results.
Benchmark performance across domains
Grok 4.7 outperformed both competitors on legal work, electrical engineering via EEBench, and the DeepSWE coding test. It scored 1,695 Elo on GDPval, which evaluates tasks performed by lawyers, nurses, and financial analysts, up from Grok 4.6's 1,605 Elo but below Fable 5.1's 1,735.
Fable 5.1 maintained clear leads on longer coding and terminal benchmarks. On Terminal-Bench 4.0, Fable scored 57.9% versus Grok 4.7's 38.0%. Fable also led on CursorBench and HealthBench Professional clinical reasoning tests, where GPT-5.6 Sol also outperformed Grok.
Grok 4.7 beat GPT-5.6 Sol on five of seven benchmarks tested.
Technical specifications
Grok 4.7 runs on 2.1 trillion parameters, 40% more than the 1.5 trillion powering Grok 4.6. xAI included supplemental SpaceX training data, including Starlink satellite telemetry and manufacturing records. The company stated the model is more likely to spend extra time on difficult problems and verify its own answers compared to Grok 4.6.


