Twenty AI models completed the same 13 knowledge-work tasks ten times each. Every output was graded against a task-specific rubric; costs use standard no-cache list prices.
Balanced Value combines quality, reliability and cost. The other lenses let you change that trade-off.
It gives most of the weight to quality, then accounts for cost and reliability. Consistency gates the cost term, so an inexpensive model with a failure tail cannot rise on price alone.
The four cards rank the same models with different priorities. Open a card to see its formula, what it rewards and where the trade-offs change the result.
Balanced Value is the default view. Switch tabs to isolate quality, price or first-attempt success, or set your own weights below. The full list scrolls while its controls stay in place.
Kimi K2.7 Code (73.1), GPT-5.5 (71.5), Claude Opus 5 (69.3) lead Balanced Value. Claude Opus 5 leads Quality Core at 87.5. All 20 rows now use the same blind full-packet review standard.
G4 OS puts these models in one workspace and lets teams route work by quality, speed and cost. Download G4 OS → · Read the benchmark code and aggregate data on GitHub.
Switch the metric tab to re-rank · Balanced is the recommended default for everyday selection.
Quality contributes 72% of Balanced Value, cost contributes 23% and reliability contributes 5%. Consistency gates the cost term. All models use the same fixed normalization scale.
Canonical score review: all 2,600 official runs were judged blindly with one frozen packet schema: full task prompt, Definition of Done, bounded source/tool evidence, and final answer. Explicit model, run, and session identity was removed from judge-facing packets.
Sonnet 5 price correction: this page uses the standard $3 input / $15 output per MTok list price. It does not use the temporary $2 / $10 launch promotion that ends August 31, 2026.
The panels use the same row order, so score range and cost can be read together. Scroll either panel to move both. The open marker is the mean; the filled marker is the median.
Agentic work costs more than tokens. Extra steps and longer tool loops add latency, orchestration load, error surface, and user-perceived unpredictability. Drag is a real cost separable from quality and price.
The two cost-aware lenses point at different defaults on purpose. Acceptable rewards cheap reliability for volume; Excellent rewards premium quality that isn't wildly overpriced. The crossover is the trade-off.
The quality leaders, the value leaders and the most consistent models are not always the same rows.
The full 2,600-run ranking has been recomputed from a uniform blind review. Pricing, latency, steps, and errors remain measured independently; the quality score no longer mixes adjudication formats.
Each cell is the mean of 10 canonical blind-reviewed runs. Switch modes to reorder the rows by raw score or score per dollar; the header remains visible while the matrix scrolls. The 2,600-run set uses one review packet and rubric across all 20 models.
"Best overall" uses task-local Quality Core. "Best cost-adjusted" uses task-local Acceptable Value. The fact that these two columns disagree on most tasks is the whole argument for routing.
Click any task to expand the full prompt, definition of done, sources and Pass@1.
Pick the task, up to three models, and either how often you run it or a plain number of runs. Totals use the median cost per run observed in the benchmark; tags compare your picks on the benchmark scores.
Inspect the methodology and aggregate data, or try the same models in G4 OS.
Tasks, prompts, adjudication rules, the canonical 2,600-run aggregate snapshot, and the historical lineage ledger live in the benchmark artifacts. Audit the methodology, reproduce the numbers, or open a PR to add a model.
View on GitHub G4 OSG4 OS gives teams access to the models on this page and supports routing work by quality, speed and cost.
Download G4 OS