agent-bench / 2026-08
14–15 Aug 2026 5 models · 12 tasks · 60 runs winner: —

Which agent for which job.

#
Model
Equal
Arman-wt
Two ways to add it up

Combined scores

The instrument

Every model, every task

0–10 per category. Darker is stronger. Mechanical tasks are scored by executable graders; rubric tasks are blind-scored by Fable and Sol, reconciled. Hover a cell for the note.

0
10 — cell shade encodes score magnitude (sequential)
Five contestants, five fingerprints

Model by model

Each radar traces one model across all 12 categories — the shape is its signature. Latency is median wall-clock per task; reliability is the share of tasks that returned a usable answer first try.

The point of the whole thing

Which model for which job

The best pick per category, with the score it earned. Where the top two are within a point I say so — a single run can't split hairs.

How to trust this

Method & honest caveats

Pinned