Agents Directory
SkillsRankingsAgents
CategoriesModelsBenchmarksCompareAgent LeaderboardSkillsRankingsAgentsAbout
/Agents Directory Blog
/Grok 4.5 benchmarks: frontier coding at a lower price
  • Artificial Analysis Intelligence IndexArtificial Analysis Intelligence Index
  • Coding benchmarks reported by SpaceXAICoding benchmarks reported by SpaceXAI
  • Efficiency and priceEfficiency and price
  • AvailabilityAvailability
  • Bottom lineBottom line
  • SourcesSources

Grok 4.5 benchmarks: frontier coding at a lower price

Grok 4.5 reaches 54 on the Artificial Analysis Intelligence Index and stays close to the leading coding models while using fewer tokens.

Jul 8, 2026•3 min read•Written byAgents Directory's profileAgents Directory@agentsdir

xxAI launched GGrok 4.5 on July 8. SpaceXAI describes it as its smartest model for coding, agentic tasks, and knowledge work. The API costs $2 per million input tokens and $6 per million output tokens, with low, medium, and high reasoning settings. High is the default.

The headline is not that Grok wins every benchmark. It does not. The interesting result is how close it gets to the leading coding models at a much lower token price and with unusually efficient output.

Artificial Analysis Intelligence Index

Artificial Analysis logoArtificial Analysis independently measured 54 for Grok 4.5 at high reasoning effort. That placed it fourth when the result was published, behind Claude Fable 5, GPT-5.5, and Claude Opus 4.8. It improved 16 points over Grok 4.3.

Artificial Analysis also measured 86.7 output tokens per second and a 500K-token context window. Those figures make Grok 4.5 a practical frontier model, not only a strong benchmark entry.

Full interactive leaderboard on our Intelligence Index page.

The chart above reads from the live catalog. It can change as benchmark publishers revise scores or add model configurations.

Coding benchmarks reported by SpaceXAI

SpaceXAI published four coding comparisons in the launch announcement. These are vendor-reported results, so the harness details matter. DeepSWE 1.0 uses each provider's own harness, while DeepSWE 1.1 uses the same mini-swe-agent harness run by DataCurve.

Grok 4.5Claude Fable 5 maxGPT-5.5 xhighClaude Opus 4.8 maxGLM-5.2
Coding
DeepSWE 1.0, provider harnesses
62.0%66.1%64.31%55.75%—
Coding
DeepSWE 1.1, DataCurve harness
53%70%67%59%44%
Terminal use
Terminal-Bench 2.1
83.3%84.3%83.4%78.9%—
Software engineering
SWE-Bench Pro
64.7%80.4%58.6%69.2%62.1%

Figures are reported by SpaceXAI in the Grok 4.5 announcement. DeepSWE 1.0 results use provider-specific harnesses and should not be read as a controlled model-only comparison.

The closest result is Terminal-Bench 2.1, where Grok 4.5's 83.3% sits within 0.1 point of GPT-5.5 and 1 point of Fable 5. On SWE-Bench Pro, it trails Fable 5 and Opus 4.8 but finishes ahead of GPT-5.5 and GLM-5.2 in SpaceXAI's table.

Artificial Analysis offers a second view of agentic coding. It scored Grok 4.5 in Grok Build at 76 on its Coding Agent Index, level with GPT-5.5 xhigh in Codex and below Fable 5 max in Claude Code. That result supports the launch positioning without treating vendor harnesses as interchangeable.

Efficiency and price

SpaceXAI reports that Grok 4.5 used 15,954 output tokens per SWE-Bench Pro task, compared with 67,020 for Opus 4.8 max. That is about 4.2 times fewer output tokens. Since Grok 4.5 also costs $2 in and $6 out per million tokens, the gap can translate into a large cost difference during long agent runs.

Token efficiency does not guarantee a cheaper completed task in every harness. Retries, tool calls, prompt caching, and failure rates all matter. Still, the combination of a 54 Intelligence Index score, near-leading agentic coding results, and $2/$6 pricing gives Grok 4.5 a credible position between premium flagships and cheaper workhorse models.

Availability

Grok 4.5 is available through the SpaceXAI API as grok-4.5, in Grok Build, and in Cursor. The launch announcement says it is temporarily free in Grok Build and Cursor. SpaceXAI's documentation also lists OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic as model gateways.

At launch, SpaceXAI said the model was not yet available through its products or API console in the European Union, with EU availability expected in mid-July.

Bottom line

Grok 4.5 is not the outright benchmark leader. Fable 5 remains ahead on the strongest coding results in SpaceXAI's own table, and GPT-5.6 Sol has since raised the frontier on other agentic evaluations. Grok 4.5's case is efficiency: frontier-adjacent coding and knowledge-work performance at $2/$6 token pricing, with output speed around 87 tokens per second.

For teams that want a capable coding model without flagship-level spend, Grok 4.5 is now one of the strongest options to test. Current scores, pricing, and availability are on the Grok 4.5 model page.

Sources

  • SpaceXAI: Introducing Grok 4.5 for launch details, vendor-reported coding benchmarks, token efficiency, and availability
  • SpaceXAI API documentation for Grok 4.5 for model ID, pricing, reasoning settings, and supported products
  • Artificial Analysis: Grok 4.5 model results for the independent Intelligence Index, speed, context, and price measurements
  • Artificial Analysis: Grok 4.5 reaches the intelligence frontier for the Coding Agent Index comparison
  • Intelligence Index on Agents Directory for the live leaderboard we keep updated
Share:
Browse:SkillsRankingsModelsBenchmarksProvidersAgentsAgent LeaderboardCompareCategories
Quick Links:AboutBlog

© 2026 Agents Directory