Independent editorial field guide · schemas 0.3.0 + 0.2.0
Choose the work.Then choose the model.
A task-first reading of the GPT-5.6 family—built to give you a defensible starting point in three minutes and enough evidence for a thirty-minute review.
Benchmarks are useful only after they become a story about capability, task, and boundary.
01 · Quick decision
Set the work. Tune the compromise.
Choose a task, decide how much headroom you want, then tell the guide whether dollars or waiting time matter more. The recommendation updates from measured score, cost, latency, and output-token data.
What are you doing?
Step 1Recommendation tendency
Step 2Money versus waiting
50 / 50Active resource metric
Step 402 · General capability overview
A broad signal before the specialist evidence.
Artificial Analysis Intelligence Index v4.1 combines nine evaluations spanning reasoning, knowledge, mathematics, programming, and increasingly agentic work. Read it as a versioned composite index—not a percentage of intelligence or a substitute for task-specific evaluation.
Artificial Analysis Intelligence Index v4.1
The benchmark is defined and maintained by Artificial Analysis. This guide uses a pre-Build-Week schema 0.2.0 snapshot whose recorded provenance and source-asset labels are preserved. It is not a live feed, not an official Artificial Analysis dataset release, and not endorsement of this project.
03 · Model identities
Four identities. No universal winner.
Official positioning describes the family; this guide adds a practical reading. Model size and reasoning effort are separate decisions, and a smaller model at a higher effort can still be the lower-resource choice.
Luna
The fastest, most affordable tier. Start here when iteration volume and resource restraint matter more than maximum headroom.
- Best first look
- Repeatable, well-bounded work
- Watch for
- Hard tasks that need sustained headroom
Terra
The balanced tier for everyday work. A strong default when Luna feels tight but Sol’s added spend is not yet justified.
- Best first look
- Mixed everyday knowledge and agent work
- Watch for
- Frontier cases where a small score gap matters
Sol
The flagship tier for the hardest coding, research, professional, cyber, and science work. Its lower effort levels can be a better balance than pushing a smaller tier to the ceiling.
- Best first look
- Complex work with expensive failure modes
- Watch for
- Diminishing returns at xhigh and max
- Guide habit
- Compare medium, high, and max before upgrading
Sol Ultra
The highest-capability setting, coordinating parallel agents for demanding work. Treat it as a deliberate escalation—not a default badge of quality.
- Best first look
- Parallelizable work where stronger results justify much higher token use
- Watch for
- Spend, coordination overhead, and tasks that do not parallelize cleanly
Reasoning is a dial, not a certificate.
Higher effort can improve scores, but it can also plateau or regress. The atlas shows the measured curve instead of assuming that each step is better.
04 · Benchmark Atlas
Capabilities first. Evidence second.
Choose a capability, then inspect the benchmark that supports it. Only one detailed benchmark is active at a time, so the page stays readable without hiding the complete configuration set.
05 · Score-only evidence
Important capability, incomplete resource picture.
These official score-table results extend the guide into specialist work. Cost, latency, output tokens, and reasoning effort were not published for these rows, so they are shown as score evidence—not forced into a resource curve.
06 · Method, limits, sources
Show the arithmetic. Keep the uncertainty.
The recommendation system is deliberately simple enough to audit. It normalizes only within the active benchmark, never compares unlike benchmark scores, and never treats project-derived resource cost as an official metric.
One global state, four resource views.
Score, cost, latency, and output tokens are normalized only within the active benchmark. Cost and tokens use log normalization; latency uses linear min–max normalization. The active metric supplies the recommendation’s resource penalty. The composite view combines all three published resource dimensions.
utility = q × scorenormalized + (1 − q) × (1 − active resourcenormalized)
w is the money-versus-waiting slider. q is 0.42, 0.63, or 0.82 for Conserve Resources, Balanced, or Quality First. The task, tendency, slider, and active metric are shared across every linked control set. All recommendations are project interpretation.
Efficient is not the same as recommended.
A point is Pareto-efficient in the active chart when no other point has at least its score at no greater active resource value, with one strict improvement.
“Fully dominated” is stricter: another point must be no worse in score, cost, latency, and output tokens, and strictly better in at least one. The recommendation star applies the project utility formula; it is not an objective winner.
Missing means missing.
- No interpolation between reasoning levels.
- No zero-fill for absent cost, latency, or tokens.
- No averages across source conflicts.
- No resource recommendation from score-only rows.
A benchmark is a controlled signal, not your production environment.
Scores may not transfer to your prompts, tools, repositories, risk tolerance, or evaluation harness. Estimated benchmark-run cost is not the same as the price of one request. Latency can depend on harness design, parallelism, and infrastructure. “Ultra” and multi-agent results are not directly interchangeable with single-agent operation. Use this guide to choose what to test next—not to avoid testing.
Normalize locally, then describe breadth.
For each capability family, the guide normalizes every score inside its own benchmark. It then averages those normalized positions by model; it never averages raw percentages, Elo ratings, F1 values, or index scores from unlike evaluations. Preference fit is averaged only across benchmarks with complete resource data. Coverage shows how many family benchmarks contribute; confidence is a project label based on evidence breadth and coverage, not statistical confidence.
Primary public source for model positioning, benchmark charts, score tables, availability, and benchmark context.
Read official source ↗Schemas 0.3.0 and 0.2.0. The page embeds 202 complete-atlas points, 58 score-only points, and 18 Artificial Analysis overview points, including missing states, notes, execution metadata, and provenance.
278 points · local · offlinePrimary benchmark-owner source for the Intelligence Index v4.1 composition, versioning, and limitations.
Read methodology ↗Capability labels, task mappings, within-benchmark synthesis, confidence labels, active-metric utility, balance points, and recommendations are original project interpretation.
unofficial