Leaderboard
Most cost-effective models
Ratings come from pairwise human votes via Bradley-Terry (Elo-style for v0), shown with a ±confidence interval — the honesty check: a model with five votes never reads as equal to one with five hundred. Rows rank by Rating; the efficiency story lives in the green Rating / ¢ column — rating earned per cent of generation cost — and the ⚡ BEST /¢ chip marks its winner. Testers screen every build — the Playable column is the share of votes that say a build actually runs and plays; ✗-heavy models sink.
| # | Model | Playable | Rating / ¢ |
|---|---|---|---|
| ★ 1 | DeepSeek V3671B params | 50%· 2/5 ✓ | ⚡ BEST /¢2976 |
| 2 | Claude Opus 4.8 | 75%· 3/9 ✓ | 96 |
| 3 | GPT-5.5 | 50%· 2/9 ✓ | 81 |
| 4 | Gemini 3.1 Pro | 75%· 3/9 ✓ | 202 |
| 5 | GLM 5.2 | 50%· 1/9 ✓ | 822 |
This board sharpens one screened build at a time. Play an unvetted build, call it playable or not — your vote lands here.
Screen a build →