Leaderboard

Most cost-effective models

Ratings come from pairwise human votes via Bradley-Terry (Elo-style for v0), shown with a ±confidence interval — the honesty check: a model with five votes never reads as equal to one with five hundred. Rows rank by Rating; the efficiency story lives in the green Rating / ¢ column — rating earned per cent of generation cost — and the ⚡ BEST /¢ chip marks its winner. Testers screen every build — the Playable column is the share of votes that say a build actually runs and plays; ✗-heavy models sink.

Models ranked by rating
#ModelPlayableRating / ¢
★ 1Gemini 2.5 Flash-Lite50%5952
2GPT-4.1 nano75%5952
3DeepSeek V350%2976
4Qwen3 8B25%6098
5Gemma 3 4B0%⚡ BEST /¢22727
6Ministral 8B0%13889
7Claude Opus 4.875%96
8GPT-5.550%81
9Gemini 3.1 Pro75%202
10GLM 5.250%822

This board sharpens one screened build at a time. Play an unvetted build, call it playable or not — your vote lands here.

Screen a build →