Updated July 30, 2026

Pick the right model for the work,
not the leaderboard.

Twenty AI models completed the same 13 knowledge-work tasks ten times each. Every output was graded against a task-specific rubric; costs use standard no-cache list prices.

Balanced Value combines quality, reliability and cost. The other lenses let you change that trade-off.

20AI models tested
13Knowledge-work tasks
10×Runs per task
2,600Total scored runs
View the ranking
The metric

Balanced Value is the practical starting point.

It gives most of the weight to quality, then accounts for cost and reliability. Consistency gates the cost term, so an inexpensive model with a failure tail cannot rise on price alone.

Read the methodology
Balanced Value=
0.72·QQuality
0.23·CCost gated by S
0.05·EReliability
Higher = better. The full breakdown lives in the full methodology.
The frontier of quality × cost

Cost and score, on the same chart.

The dashed line marks the efficient frontier. Labels show the current leaders and frontier; hover or tap any bubble for details.
Four lenses at a glance

Choose the lens that matches the job.

The four cards rank the same models with different priorities. Open a card to see its formula, what it rewards and where the trade-offs change the result.

Final Ranking

Rank the models by what matters to you.

Balanced Value is the default view. Switch tabs to isolate quality, price or first-attempt success, or set your own weights below. The full list scrolls while its controls stay in place.

Your priorities
Quality72%
Cost23%
Reliability5%
Drag to re-rank the board with your own weights. The default 72 / 23 / 5 mirrors Balanced Value; cost stays gated by consistency.
# Model Balanced (higher = better) 73.1
The decision
73.1
top Balanced Value
Kimi K2.7 Code
Balanced #1
GPT-5.5
Balanced #2
Opus 5
Balanced #3

Kimi K2.7 Code leads the canonical Balanced view.

Kimi K2.7 Code (73.1), GPT-5.5 (71.5), Claude Opus 5 (69.3) lead Balanced Value. Claude Opus 5 leads Quality Core at 87.5. All 20 rows now use the same blind full-packet review standard.

Balanced leaders
Kimi K2.7 CodeGPT-5.5Opus 5
The current cost-aware leaders after canonical review of all 2,600 official runs.
Quality leaders
Opus 5GPT-5.5GPT-5.6 Sol
Claude Opus 5 leads Quality Core at 87.5. Cost is excluded from this lens.
One review standard
20 models13 tasks10 runs
Every run was judged from the full task prompt, Definition of Done, bounded tool evidence, and final answer with explicit model identity removed.

G4 OS puts these models in one workspace and lets teams route work by quality, speed and cost. Download G4 OS →  ·  Read the benchmark code and aggregate data on GitHub.

Switch the metric tab to re-rank · Balanced is the recommended default for everyday selection.

Methodology

How the scores are calculated.

Quality contributes 72% of Balanced Value, cost contributes 23% and reliability contributes 5%. Consistency gates the cost term. All models use the same fixed normalization scale.

Canonical score review: all 2,600 official runs were judged blindly with one frozen packet schema: full task prompt, Definition of Done, bounded source/tool evidence, and final answer. Explicit model, run, and session identity was removed from judge-facing packets.

Sonnet 5 price correction: this page uses the standard $3 input / $15 output per MTok list price. It does not use the temporary $2 / $10 launch promotion that ends August 31, 2026.

Drill-down

How the scores spread, model by model.

The panels use the same row order, so score range and cost can be read together. Scroll either panel to move both. The open marker is the mean; the filled marker is the median.

Quality score per run (0–100)

ordered by reliability-adjusted score (μ − σ/2) · mean (○) · median (●) · min → max

Cost per run (USD)

no-cache · same row order as quality — read across to see cost vs reliability

Operational Drag: wasted steps and time

Agentic work costs more than tokens. Extra steps and longer tool loops add latency, orchestration load, error surface, and user-perceived unpredictability. Drag is a real cost separable from quality and price.

1.7Lowest · Fable 5
91.2Highest · Gemini Flash

Acceptable vs Excellent: two lenses, two defaults

The two cost-aware lenses point at different defaults on purpose. Acceptable rewards cheap reliability for volume; Excellent rewards premium quality that isn't wildly overpriced. The crossover is the trade-off.

Why balanced wins

What changes when cost and variance enter the ranking.

The quality leaders, the value leaders and the most consistent models are not always the same rows.

Quality is crowded
65.1–87.5
Claude Opus 5 (87.5), GPT-5.5 (85.5), GPT-5.6 Sol (84.2) lead the canonical Quality Core view.
Cost is not
239×
$0.011 (DeepSeek) to $2.63 (Opus 5) per no-cache run. The 239× cost spread dwarfs the quality spread. At scale, cost is the variable that moves.
Downside tails
17.7%
4 canonical model rows show zero catastrophic runs; the highest observed catastrophic rate is 17.7%.
Different #1 per lens
2 winners
Kimi K2.7 Code leads Balanced; Claude Opus 5 leads Excellent Value.
Canonical recalibration

One standard across every model.

The full 2,600-run ranking has been recomputed from a uniform blind review. Pricing, latency, steps, and errors remain measured independently; the quality score no longer mixes adjudication formats.

Per-task performance

Model by model, task by task.

Each cell is the mean of 10 canonical blind-reviewed runs. Switch modes to reorder the rows by raw score or score per dollar; the header remains visible while the matrix scrolls. The 2,600-run set uses one review packet and rubric across all 20 models.

Mean score on the task · pure quality signal
Low High Color scales within each task column · border marks the column leader
Task winners

A different #1 on almost every task.

"Best overall" uses task-local Quality Core. "Best cost-adjusted" uses task-local Acceptable Value. The fact that these two columns disagree on most tasks is the whole argument for routing.

The 13 tasks

Inside the thirteen tasks.

Click any task to expand the full prompt, definition of done, sources and Pass@1.

What would this task cost you?

Pick the task, up to three models, and either how often you run it or a plain number of runs. Totals use the median cost per run observed in the benchmark; tags compare your picks on the benchmark scores.

More

Take it further.

Inspect the methodology and aggregate data, or try the same models in G4 OS.