"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok · TryAI
Added
Article: Positive

Researchers tested four AI models in a digital drawing arena to evaluate their ability to use tools and follow creative prompts. GPT-5.6 Sol proved to be the most effective artist, outperforming the much more expensive Claude Fable 5 and the underperforming Grok 4.5. The study also found that AI models often over-edit their work, leading to a decline in quality after an initial peak.
Key Points
- GPT-5.6 Sol was the top performer, delivering the best artistic quality and detail at a very low relative cost.
- Claude Fable 5 was the least efficient model, costing roughly 20 times more than its competitors while producing inferior results.
- Models tend to plateau early and actually decrease the quality of their work if they continue editing past their peak performance.
- Grok 4.5 struggled significantly with the visual tasks, highlighting a gap in its current visual reasoning capabilities.
- Gemini 3.6 Flash achieved high structural similarity scores and was the most frequent self-reviewer, though its subjective quality trailed Sol.