Leaderboard

Most cost-effective models

Ratings come from pairwise human votes via Bradley-Terry (Elo-style for v0), shown with a ±confidence interval — the honesty check: a model with five votes never reads as equal to one with five hundred. Rows rank by Rating; the efficiency story lives in the green Rating / ¢ column — rating earned per cent of generation cost — and the ⚡ BEST /¢ chip marks its winner. Testers screen every build — the Playable column is the share of votes that say a build actually runs and plays; ✗-heavy models sink.

Models ranked by rating
#ModelPlayableRating / ¢
★ 1DeepSeek V350%⚡ BEST /¢2976
2Claude Opus 4.875%96
3GPT-5.550%81
4Gemini 3.1 Pro75%202
5GLM 5.250%822

This board sharpens one screened build at a time. Play an unvetted build, call it playable or not — your vote lands here.

Screen a build →