0–10 per category. Darker is stronger. Mechanical tasks are scored by executable graders; rubric tasks are blind-scored by Fable and Sol, reconciled. Hover a cell for the note.
Each radar traces one model across all 12 categories — the shape is its signature. Latency is median wall-clock per task; reliability is the share of tasks that returned a usable answer first try.
The best pick per category, with the score it earned. Where the top two are within a point I say so — a single run can't split hairs.