Independent editorial field guide · schemas 0.3.0 + 0.2.0

Choose the work.Then choose the model.

A task-first reading of the GPT-5.6 family—built to give you a defensible starting point in three minutes and enough evidence for a thirty-minute review.

00
Benchmarks are useful only after they become a story about capability, task, and boundary.

01 · Quick decision

Set the work. Tune the compromise.

Choose a task, decide how much headroom you want, then tell the guide whether dollars or waiting time matter more. The recommendation updates from measured score, cost, latency, and output-token data.

What are you doing?

Step 1

Recommendation tendency

Step 2

Money versus waiting

50 / 50
Save more moneySave more waiting time

This changes the resource penalty, not how long the model is allowed to reason. Output tokens retain a fixed 20% share.

Active resource metric

Step 4
Project interpretation · not official advice

Solmedium reasoning

02 · General capability overview

A broad signal before the specialist evidence.

Artificial Analysis Intelligence Index v4.1 combines nine evaluations spanning reasoning, knowledge, mathematics, programming, and increasingly agentic work. Read it as a versioned composite index—not a percentage of intelligence or a substitute for task-specific evaluation.

Published index · project interpretation layer

Artificial Analysis Intelligence Index v4.1

The benchmark is defined and maintained by Artificial Analysis. This guide uses a pre-Build-Week schema 0.2.0 snapshot whose recorded provenance and source-asset labels are preserved. It is not a live feed, not an official Artificial Analysis dataset release, and not endorsement of this project.

Evidence18 complete points
ModelsLuna · Terra · Sol
Score typeVersioned composite index
Snapshot schema0.2.0 · generatedAt null
Capability representedBroad cross-domain intelligence across nine constituent evaluations.
Tasks it resemblesA mix of agentic work, coding, scientific reasoning, knowledge, mathematics, and general problem solving.
What it does not measureYour exact workflow, safety, reliability, taste, deployment fit, or a universal percentage of intelligence.

03 · Model identities

Four identities. No universal winner.

Official positioning describes the family; this guide adds a practical reading. Model size and reasoning effort are separate decisions, and a smaller model at a higher effort can still be the lower-resource choice.

Luna

The fastest, most affordable tier. Start here when iteration volume and resource restraint matter more than maximum headroom.

Best first look
Repeatable, well-bounded work
Watch for
Hard tasks that need sustained headroom

Terra

The balanced tier for everyday work. A strong default when Luna feels tight but Sol’s added spend is not yet justified.

Best first look
Mixed everyday knowledge and agent work
Watch for
Frontier cases where a small score gap matters

Sol

The flagship tier for the hardest coding, research, professional, cyber, and science work. Its lower effort levels can be a better balance than pushing a smaller tier to the ceiling.

Best first look
Complex work with expensive failure modes
Watch for
Diminishing returns at xhigh and max
Guide habit
Compare medium, high, and max before upgrading

Sol Ultra

The highest-capability setting, coordinating parallel agents for demanding work. Treat it as a deliberate escalation—not a default badge of quality.

Best first look
Parallelizable work where stronger results justify much higher token use
Watch for
Spend, coordination overhead, and tasks that do not parallelize cleanly

Reasoning is a dial, not a certificate.

Higher effort can improve scores, but it can also plateau or regress. The atlas shows the measured curve instead of assuming that each step is better.

noneReported only where the source exposes no explicit reasoning level.
lowFast exploration and bounded work.
mediumPractical default for many hard tasks.
highMore search, checking, and revision.
xhighSpecialized escalation with rising resource use.
maxLargest measured effort; inspect marginal gain first.

04 · Benchmark Atlas

Capabilities first. Evidence second.

Choose a capability, then inspect the benchmark that supports it. Only one detailed benchmark is active at a time, so the page stays readable without hiding the complete configuration set.

7capability families
19benchmark definitions
178points with cost, latency, and tokens
24points with explicit missing-resource states

05 · Score-only evidence

Important capability, incomplete resource picture.

These official score-table results extend the guide into specialist work. Cost, latency, output tokens, and reasoning effort were not published for these rows, so they are shown as score evidence—not forced into a resource curve.

06 · Method, limits, sources

Show the arithmetic. Keep the uncertainty.

The recommendation system is deliberately simple enough to audit. It normalizes only within the active benchmark, never compares unlike benchmark scores, and never treats project-derived resource cost as an official metric.

Project-specific preference fit

One global state, four resource views.

Score, cost, latency, and output tokens are normalized only within the active benchmark. Cost and tokens use log normalization; latency uses linear min–max normalization. The active metric supplies the recommendation’s resource penalty. The composite view combines all three published resource dimensions.

composite = 0.8 × [(1 − w) × costlog + w × latencylinear] + 0.2 × tokenslog
utility = q × scorenormalized + (1 − q) × (1 − active resourcenormalized)

w is the money-versus-waiting slider. q is 0.42, 0.63, or 0.82 for Conserve Resources, Balanced, or Quality First. The task, tendency, slider, and active metric are shared across every linked control set. All recommendations are project interpretation.

Pareto and dominance

Efficient is not the same as recommended.

A point is Pareto-efficient in the active chart when no other point has at least its score at no greater active resource value, with one strict improvement.

“Fully dominated” is stricter: another point must be no worse in score, cost, latency, and output tokens, and strictly better in at least one. The recommendation star applies the project utility formula; it is not an objective winner.

Missing-data policy

Missing means missing.

  • No interpolation between reasoning levels.
  • No zero-fill for absent cost, latency, or tokens.
  • No averages across source conflicts.
  • No resource recommendation from score-only rows.
Limitations

A benchmark is a controlled signal, not your production environment.

Scores may not transfer to your prompts, tools, repositories, risk tolerance, or evaluation harness. Estimated benchmark-run cost is not the same as the price of one request. Latency can depend on harness design, parallelism, and infrastructure. “Ultra” and multi-agent results are not directly interchangeable with single-agent operation. Use this guide to choose what to test next—not to avoid testing.

Capability synthesis

Normalize locally, then describe breadth.

For each capability family, the guide normalizes every score inside its own benchmark. It then averages those normalized positions by model; it never averages raw percentages, Elo ratings, F1 values, or index scores from unlike evaluations. Preference fit is averaged only across benchmarks with complete resource data. Coverage shows how many family benchmarks contribute; confidence is a project label based on evidence breadth and coverage, not statistical confidence.

AOpenAI GPT-5.6 release

Primary public source for model positioning, benchmark charts, score tables, availability, and benchmark context.

Read official source ↗
BEmbedded benchmark datasets

Schemas 0.3.0 and 0.2.0. The page embeds 202 complete-atlas points, 58 score-only points, and 18 Artificial Analysis overview points, including missing states, notes, execution metadata, and provenance.

278 points · local · offline
CArtificial Analysis methodology

Primary benchmark-owner source for the Intelligence Index v4.1 composition, versioning, and limitations.

Read methodology ↗
DIndependent interpretation layer

Capability labels, task mappings, within-benchmark synthesis, confidence labels, active-metric utility, balance points, and recommendations are original project interpretation.

unofficial