GPT-5.6 Sol, Terra and Luna benchmark results
GPT-5.6 Sol leads OpenAI's new family, while Terra and Luna target lower costs. Here are the independent and official benchmark results that matter.
•5 min read•Written by
Agents Directory@agentsdir
OpenAI released the GPT-5.6 family for general availability on July 9, 2026. There is no single model called "GPT-5.6 Sol Terra." The family has three separate tiers: GPT-5.6 Sol is the flagship, GPT-5.6 Terra balances capability and cost, and GPT-5.6 Luna is the fastest and cheapest option.
The short version of the GPT-5.6 benchmarks is that Sol is OpenAI's strongest all-round model, especially for terminal work, computer use, and cybersecurity. Terra stays surprisingly close for half the token price. Luna gives up more capability, but it remains competitive with the previous flagship on several professional and coding tests. This article separates independently maintained benchmark data from OpenAI's own launch evaluations.
The GPT-5.6 family at a glance
All three models launched across ChatGPT, Codex, and the OpenAI API. OpenAI also introduced max reasoning for longer single-agent work and ultra, a multi-agent mode that coordinates four agents by default. The three API tiers use the same generation number but have different prices and performance targets:
- GPT-5.6 Sol: $5 per million input tokens and $30 per million output tokens. This is the flagship for difficult coding, research, computer use, and professional work.
- GPT-5.6 Terra: $2.50 input and $15 output. It costs half as much as Sol and is positioned as the balanced everyday model.
- GPT-5.6 Luna: $1 input and $6 output. It is the fastest and most cost-efficient member of the family.
These are durable capability tiers, not reasoning settings. A user can choose Sol, Terra, or Luna, then select an available reasoning effort for that model.
Independent GPT-5.6 benchmark: Artificial Analysis
Artificial Analysis maintains an independent composite intelligence index spanning reasoning, coding, mathematics, scientific questions, and agentic tasks. The chart below reads from the benchmark data tracked by Agents Directory, so it can change as Artificial Analysis reruns models or updates its methodology.
OpenAI's launch table cited Artificial Analysis Intelligence Index v4.1 scores of 58.9 for Sol, 55.0 for Terra, and 51.2 for Luna. In that same table, Claude Fable 5 scored 59.9, GPT-5.5 scored 54.8, and Claude Opus 4.8 scored 55.7. That makes Terra the interesting value result: it slightly clears GPT-5.5 on the index at half the input and output price. The live chart above is the independent record to follow because its values can move after launch day.
OpenAI's reported benchmark results
The following results come from OpenAI's GPT-5.6 launch post, not from an Agents Directory rerun. They are useful for comparing the three tiers under one published test setup, but they remain vendor-reported numbers. Harnesses, reasoning effort, tool access, and token budgets differ by evaluation, so the table should not be read as a universal model ranking.
| GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-5.5 | Claude Fable 5 | Claude Opus 4.8 | |
|---|---|---|---|---|---|---|
Professional work Agents' Last Exam | 52.7% | 50.4% | 50.3% | 46.9% | 40.5% | 45.2% |
General intelligence Artificial Analysis Intelligence Index v4.1 | 58.9 | 55.0 | 51.2 | 54.8 | 59.9 | 55.7 |
Coding Artificial Analysis Coding Agent Index v1.1 | 80.0 | 77.4 | 74.6 | 76.4 | 77.2 | 72.5 |
Agentic coding SWE-Bench Pro | 64.6% | 63.4% | 62.7% | 59.4% | 80.0% | 69.2% |
Agentic coding DeepSWE v1.1 | 72.7% | 69.6% | 67.2% | 67.0% | 69.7% | 59.0% |
Terminal work Terminal-Bench 2.1 | 88.8% 91.9% with Sol Ultra | 87.4% | 84.7% | 85.6% | 83.1% | 78.9% |
Biology GeneBench Pro | 28.7% | 23.3% | 10.8% | 12.0% | — | 16.0% |
Computer use OSWorld 2.0 | 62.6% | 50.2% | 45.6% | 47.5% | — | 54.8% |
Agentic browsing BrowseComp | 90.4% 92.2% with Sol Ultra | 87.5% | 83.3% | 84.4% | — | 84.3% |
Cybersecurity SEC-Bench Pro | 71.2% 74.3% with Sol Ultra | 57.7% | 48.9% | 45.8% | — | — |
Cybersecurity ExploitBench | 73.5% | 52.9% | 33.2% | 47.9% | — | 40.0% |
Cybersecurity ExploitGym, six-hour cap | 33.7% | 23.2% | 12.4% | 15.1% | — | — |
Source: OpenAI's July 9, 2026 GPT-5.6 launch table. A blank cell means OpenAI did not report that model in the corresponding row. Sol Ultra coordinates multiple agents, so its scores are not directly comparable with the standard single-model configurations.
What the benchmark table says
Sol has the clearest lead on tool-heavy work. Its 88.8% Terminal-Bench 2.1 score is 3.2 points above GPT-5.5, while OSWorld 2.0 rises from 47.5% to 62.6%. Sol Ultra pushes Terminal-Bench to 91.9% and BrowseComp to 92.2%, but that mode coordinates multiple agents and spends more compute. It should be treated as a separate system configuration, not a free uplift to the base model.
Terra is the price-performance story. It posts 63.4% on SWE-Bench Pro compared with GPT-5.5 at 59.4%, 69.6% on DeepSWE compared with 67.0%, and 87.5% on BrowseComp compared with 84.4%. Those are OpenAI's evaluations, but the pattern also appears in the independently maintained Artificial Analysis index. At $2.50 in and $15 out, Terra costs half as much as both Sol and GPT-5.5.
Luna is not simply a small chat model. It reaches 62.7% on SWE-Bench Pro and 67.2% on DeepSWE, both above the GPT-5.5 values in OpenAI's table. Its weaker areas are more specialized work: GeneBench Pro falls to 10.8%, and ExploitBench falls to 33.2%. At $1 in and $6 out, it is aimed at workloads where throughput and cost matter more than the highest possible pass rate.
Caveats before choosing a winner
First, most of the detailed scores above were selected and published by OpenAI. Even when the benchmark itself is external, a vendor-run result can depend on a private harness, prompt, reasoning budget, tool setup, or unreleased serving configuration. The independent Artificial Analysis chart is therefore the best live cross-check in this article.
Second, Claude Fable 5 is missing from some biology and cybersecurity rows. OpenAI says it excluded Fable from GeneBench Pro because the model refuses most advanced biology questions. A missing result is not a zero, but it also means the table cannot support a clean head-to-head conclusion for every domain.
Third, ultra is a multi-agent system. Its BrowseComp, Terminal-Bench, and SEC-Bench Pro results measure coordinated parallel work, while the standard Sol, Terra, and Luna columns measure ordinary model configurations. Ultra may finish hard tasks faster, but it can consume more tokens and should be evaluated against the cost and latency of the whole run.
Finally, benchmark strength does not guarantee the same ordering in your agent. Repository size, tool reliability, context shape, retry policy, and task mix can matter as much as the checkpoint. Run a representative internal evaluation before moving production traffic.
Verdict
GPT-5.6 Sol is the strongest choice in the family when task success matters more than token price. Its best evidence is the broad lead across OpenAI's terminal, browsing, computer-use, and cybersecurity evaluations. GPT-5.6 Terra is likely the practical default for many agents because it stays close to Sol, beats GPT-5.5 on several reported tests, and costs half as much. GPT-5.6 Luna is the throughput option, with coding results that remain strong for its $1 in and $6 out price.
For current independent placement, use the Artificial Analysis Intelligence Index leaderboard. Pricing, availability, and benchmark records are tracked on the GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna model pages.
Sources
- OpenAI: GPT-5.6, frontier intelligence that scales with your ambition for general availability, pricing, product tiers, and the reported capability table
- OpenAI: Previewing GPT-5.6 Sol for the June 26 preview, the initial evaluation framing, and the Sol, Terra, and Luna naming system
- OpenAI: GPT-5.6 system card for safety methodology, capability evaluations, and evaluation limitations
- Artificial Analysis Intelligence Index for the independent composite benchmark and methodology
- Artificial Analysis Intelligence Index on Agents Directory for the live leaderboard we keep updated