Agents Directory
SkillsRankingsAgents
CategoriesModelsBenchmarksCompareAgent LeaderboardSkillsRankingsAgentsAbout
/Agents Directory Blog
/GPT-5.6 Sol, Terra and Luna benchmark results
  • The GPT-5.6 family at a glanceThe GPT-5.6 family at a glance
  • Independent GPT-5.6 benchmark: Artificial AnalysisIndependent GPT-5.6 benchmark: Artificial Analysis
  • OpenAI's reported benchmark resultsOpenAI's reported benchmark results
  • What the benchmark table saysWhat the benchmark table says
  • Caveats before choosing a winnerCaveats before choosing a winner
  • VerdictVerdict
  • SourcesSources

GPT-5.6 Sol, Terra and Luna benchmark results

GPT-5.6 Sol leads OpenAI's new family, while Terra and Luna target lower costs. Here are the independent and official benchmark results that matter.

Jul 10, 2026•5 min read•Written byAgents Directory's profileAgents Directory@agentsdir

OpenAIOpenAI released the OpenAIGPT-5.6 family for general availability on July 9, 2026. There is no single model called "GPT-5.6 Sol Terra." The family has three separate tiers: GPT-5.6 Sol is the flagship, GPT-5.6 Terra balances capability and cost, and GPT-5.6 Luna is the fastest and cheapest option.

The short version of the GPT-5.6 benchmarks is that Sol is OpenAI's strongest all-round model, especially for terminal work, computer use, and cybersecurity. Terra stays surprisingly close for half the token price. Luna gives up more capability, but it remains competitive with the previous flagship on several professional and coding tests. This article separates independently maintained benchmark data from OpenAI's own launch evaluations.

The GPT-5.6 family at a glance

All three models launched across ChatGPT, Codex, and the OpenAI API. OpenAI also introduced max reasoning for longer single-agent work and ultra, a multi-agent mode that coordinates four agents by default. The three API tiers use the same generation number but have different prices and performance targets:

  • GPT-5.6 Sol: $5 per million input tokens and $30 per million output tokens. This is the flagship for difficult coding, research, computer use, and professional work.
  • GPT-5.6 Terra: $2.50 input and $15 output. It costs half as much as Sol and is positioned as the balanced everyday model.
  • GPT-5.6 Luna: $1 input and $6 output. It is the fastest and most cost-efficient member of the family.

These are durable capability tiers, not reasoning settings. A user can choose Sol, Terra, or Luna, then select an available reasoning effort for that model.

Independent GPT-5.6 benchmark: Artificial Analysis

Artificial Analysis logoArtificial Analysis maintains an independent composite intelligence index spanning reasoning, coding, mathematics, scientific questions, and agentic tasks. The chart below reads from the benchmark data tracked by Agents Directory, so it can change as Artificial Analysis reruns models or updates its methodology.

Full interactive leaderboard on our Artificial Analysis Intelligence Index page.

OpenAI's launch table cited Artificial Analysis Intelligence Index v4.1 scores of 58.9 for Sol, 55.0 for Terra, and 51.2 for Luna. In that same table, Claude Fable 5 scored 59.9, GPT-5.5 scored 54.8, and Claude Opus 4.8 scored 55.7. That makes Terra the interesting value result: it slightly clears GPT-5.5 on the index at half the input and output price. The live chart above is the independent record to follow because its values can move after launch day.

OpenAI's reported benchmark results

The following results come from OpenAI's GPT-5.6 launch post, not from an Agents Directory rerun. They are useful for comparing the three tiers under one published test setup, but they remain vendor-reported numbers. Harnesses, reasoning effort, tool access, and token budgets differ by evaluation, so the table should not be read as a universal model ranking.

GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaGPT-5.5Claude Fable 5Claude Opus 4.8
Professional work
Agents' Last Exam
52.7%50.4%50.3%46.9%40.5%45.2%
General intelligence
Artificial Analysis Intelligence Index v4.1
58.955.051.254.859.955.7
Coding
Artificial Analysis Coding Agent Index v1.1
80.077.474.676.477.272.5
Agentic coding
SWE-Bench Pro
64.6%63.4%62.7%59.4%80.0%69.2%
Agentic coding
DeepSWE v1.1
72.7%69.6%67.2%67.0%69.7%59.0%
Terminal work
Terminal-Bench 2.1
88.8%
91.9% with Sol Ultra
87.4%84.7%85.6%83.1%78.9%
Biology
GeneBench Pro
28.7%23.3%10.8%12.0%—16.0%
Computer use
OSWorld 2.0
62.6%50.2%45.6%47.5%—54.8%
Agentic browsing
BrowseComp
90.4%
92.2% with Sol Ultra
87.5%83.3%84.4%—84.3%
Cybersecurity
SEC-Bench Pro
71.2%
74.3% with Sol Ultra
57.7%48.9%45.8%——
Cybersecurity
ExploitBench
73.5%52.9%33.2%47.9%—40.0%
Cybersecurity
ExploitGym, six-hour cap
33.7%23.2%12.4%15.1%——

Source: OpenAI's July 9, 2026 GPT-5.6 launch table. A blank cell means OpenAI did not report that model in the corresponding row. Sol Ultra coordinates multiple agents, so its scores are not directly comparable with the standard single-model configurations.

What the benchmark table says

Sol has the clearest lead on tool-heavy work. Its 88.8% Terminal-Bench 2.1 score is 3.2 points above GPT-5.5, while OSWorld 2.0 rises from 47.5% to 62.6%. Sol Ultra pushes Terminal-Bench to 91.9% and BrowseComp to 92.2%, but that mode coordinates multiple agents and spends more compute. It should be treated as a separate system configuration, not a free uplift to the base model.

Terra is the price-performance story. It posts 63.4% on SWE-Bench Pro compared with GPT-5.5 at 59.4%, 69.6% on DeepSWE compared with 67.0%, and 87.5% on BrowseComp compared with 84.4%. Those are OpenAI's evaluations, but the pattern also appears in the independently maintained Artificial Analysis index. At $2.50 in and $15 out, Terra costs half as much as both Sol and GPT-5.5.

Luna is not simply a small chat model. It reaches 62.7% on SWE-Bench Pro and 67.2% on DeepSWE, both above the GPT-5.5 values in OpenAI's table. Its weaker areas are more specialized work: GeneBench Pro falls to 10.8%, and ExploitBench falls to 33.2%. At $1 in and $6 out, it is aimed at workloads where throughput and cost matter more than the highest possible pass rate.

Caveats before choosing a winner

First, most of the detailed scores above were selected and published by OpenAI. Even when the benchmark itself is external, a vendor-run result can depend on a private harness, prompt, reasoning budget, tool setup, or unreleased serving configuration. The independent Artificial Analysis chart is therefore the best live cross-check in this article.

Second, Claude Fable 5 is missing from some biology and cybersecurity rows. OpenAI says it excluded Fable from GeneBench Pro because the model refuses most advanced biology questions. A missing result is not a zero, but it also means the table cannot support a clean head-to-head conclusion for every domain.

Third, ultra is a multi-agent system. Its BrowseComp, Terminal-Bench, and SEC-Bench Pro results measure coordinated parallel work, while the standard Sol, Terra, and Luna columns measure ordinary model configurations. Ultra may finish hard tasks faster, but it can consume more tokens and should be evaluated against the cost and latency of the whole run.

Finally, benchmark strength does not guarantee the same ordering in your agent. Repository size, tool reliability, context shape, retry policy, and task mix can matter as much as the checkpoint. Run a representative internal evaluation before moving production traffic.

Verdict

GPT-5.6 Sol is the strongest choice in the family when task success matters more than token price. Its best evidence is the broad lead across OpenAI's terminal, browsing, computer-use, and cybersecurity evaluations. GPT-5.6 Terra is likely the practical default for many agents because it stays close to Sol, beats GPT-5.5 on several reported tests, and costs half as much. GPT-5.6 Luna is the throughput option, with coding results that remain strong for its $1 in and $6 out price.

For current independent placement, use the Artificial Analysis Intelligence Index leaderboard. Pricing, availability, and benchmark records are tracked on the GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna model pages.

Sources

  • OpenAI: GPT-5.6, frontier intelligence that scales with your ambition for general availability, pricing, product tiers, and the reported capability table
  • OpenAI: Previewing GPT-5.6 Sol for the June 26 preview, the initial evaluation framing, and the Sol, Terra, and Luna naming system
  • OpenAI: GPT-5.6 system card for safety methodology, capability evaluations, and evaluation limitations
  • Artificial Analysis Intelligence Index for the independent composite benchmark and methodology
  • Artificial Analysis Intelligence Index on Agents Directory for the live leaderboard we keep updated
Share:
Browse:SkillsRankingsModelsBenchmarksProvidersAgentsAgent LeaderboardCompareCategories
Quick Links:AboutBlog

© 2026 Agents Directory