Is Pelicanmaxxing Real? Testing AI Models for Benchmark Gaming
Article: NeutralCommunity: Very PositiveMixed
The author tested whether AI labs are specifically gaming the 'pelican on a bicycle' benchmark by comparing it to other animal-vehicle combinations. The study found that pelicans and bicycles are actually among the lowest-scoring subjects and show no signs of special optimization. Ultimately, the results suggest that models are improving their general SVG capabilities rather than memorizing specific prompts.
Key Points
- The experiment compared the famous 'pelican on a bicycle' prompt against 47 other animal-vehicle variations across seven top AI models.
- Quantitative data showed that pelicans and bicycles are actually harder for models to draw than other subjects, ranking near the bottom of the quality scale.
- Statistical regression analysis failed to find any significant 'benchmark-specific' boost for any of the tested models.
- Common compositional traits, such as the pelican facing right, were found to be typical for those subjects rather than evidence of memorization.
Sentiment
The overall sentiment is analytical and cautiously optimistic; users appreciate the rigorous attempt to debunk 'cheating' rumors while remaining skeptical of the models' underlying structural understanding and aesthetic quality.
In Agreement
- The methodology of using a grid of 48 animal-vehicle combinations is a robust way to test for benchmark gaming.
- AI labs are likely 'SVGmaxxing' (improving general vector graphics skills) rather than just 'pelicanmaxxing,' which is a legitimate and useful development.
- SVG generation is a valuable proxy for spatial reasoning and general programming ability.
- The benchmark has successfully pressured labs to fix previously poor SVG performance.
- The right-facing bias in bicycle images is likely a reflection of real-world photography conventions (showing the drivetrain) rather than specific cheating.
Opposed
- The study's reliance on an LLM judge and 'vibes' makes the results subjective and potentially statistically unsound.
- The 100% right-facing orientation for the specific pelican-bicycle prompt is strong evidence that models have been exposed to this specific data during training.
- AI-generated SVGs are still aesthetically poor ('uglier than sin') and often structurally impossible, limiting their professional utility.
- Goodhart's Law suggests that because this benchmark is now a known target, it has lost its value as a measure of general intelligence.
- Specialized models (like Recraft V4) are significantly better at SVG generation than general-purpose LLMs, making LLM SVG capabilities feel like a 'dancing dog' miracle rather than a practical tool.