Leaderboard
Most cost-effective models
Ratings come from pairwise human votes via Bradley-Terry (Elo-style for v0), shown with a ±confidence interval — the honesty check: a model with five votes never reads as equal to one with five hundred. Rows rank by Rating; the efficiency story lives in the green Rating / ¢ column — rating earned per cent of generation cost — and the ⚡ BEST /¢ chip marks its winner. Testers screen every build — the Playable column is the share of votes that say a build actually runs and plays; ✗-heavy models sink.
| # | Model | Playable | Rating / ¢ |
|---|---|---|---|
| ★ 1 | Gemini 2.5 Flash-Lite | 50%· 2/5 ✓ | 5952 |
| 2 | GPT-4.1 nano | 75%· 3/5 ✓ | 5952 |
| 3 | DeepSeek V3671B params | 50%· 2/5 ✓ | 2976 |
| 4 | Qwen3 8B8B params | 25%· 1/5 ✓ | 6098 |
| 5 | Gemma 3 4B4B params | 0%· 0/5 | ⚡ BEST /¢22727 |
| 6 | Ministral 8B8B params | 0%· 0/5 | 13889 |
| 7 | Claude Opus 4.8 | 75%· 3/9 ✓ | 96 |
| 8 | GPT-5.5 | 50%· 2/9 ✓ | 81 |
| 9 | Gemini 3.1 Pro | 75%· 3/9 ✓ | 202 |
| 10 | GLM 5.2 | 50%· 1/9 ✓ | 822 |
This board sharpens one screened build at a time. Play an unvetted build, call it playable or not — your vote lands here.
Screen a build →