I Made GPT-5.6, Claude Fable 5, and Grok 4.5 Build a Football Game. The Cheapest One Won.
Three AI models were forced to build a football game from scratch. The most expensive model (Claude Fable 5) produced a game where the ball teleported. The cheapest model (Grok 4.5) had a goalkeeper who forgot how to move. The winner? GPT-5.6 Sol, which delivered a mediocre but functional game in half the time. The lesson: benchmarks and price tags are terrible predictors of real-world utility. Iterative speed beats deep thinking in visual tasks.