AI Model Showdown: Comparing Cost and Quality Across 11 Coding Agents
Article: Very PositiveCommunity: NeutralMixed

Netlify tested 11 different AI models through their new OpenRouter partnership to compare their performance in building a simple coffee shop website. The experiment highlighted a massive range in credit costs and design quality, with high-end models like Claude Opus offering superior detail at a premium price. Ultimately, the study demonstrates that users must balance their desire for automated, high-quality design against their available credit budget.
Key Points
- Netlify's partnership with OpenRouter introduces a wide array of open and frontier models to the Agent Runners platform.
- Model performance varies significantly in terms of design intuition, content richness, and adherence to functional requirements.
- Credit costs range drastically, from over 1,000 credits for a single high-effort Claude Opus run to less than 2 credits for DeepSeek V4 Flash.
- Premium models like Claude Opus and GPT-5.6 Sol generally provide better out-of-the-box design and self-correction.
- The choice of model should be dictated by the user's budget and whether they prefer a fully automated result or an iterative, guided process.
Sentiment
Skeptical but engaged; users appreciate the data on cost and variety but remain critical of the benchmarking methodology and its relevance to professional workflows.
In Agreement
- The inclusion of actual credit/token costs alongside model output is a highly useful metric.
- Observing the distinct 'personalities' and design sensibilities of different models is valuable for understanding their training biases.
- The study correctly identifies a massive disparity in efficiency and token usage between frontier and open-weights models.
- Task-specific evaluations are generally more interesting and useful than generic, abstract benchmarks.
Opposed
- One-shot prompts are unrepresentative of real-world development and are only useful for 'vibe coders' or non-programmers.
- A sample size of three runs is statistically insufficient to account for the inherent randomness and variance of LLM outputs.
- Popular benchmarks are often contaminated because their solutions eventually enter the training datasets of newer model iterations.
- Premium models often produce 'AI-vibe' designs that prioritize aesthetic complexity over basic user utility, such as easily finding a menu or address.