My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
Posted by thebigship 17 hours ago
Comments
Comment by jnwatson 11 hours ago
Definitely has some creative flourishes.
(I made no extra prompting. Just the above text. Single shot.)
Comment by xeromal 8 hours ago
Comment by zombot 8 hours ago
Comment by amarant 7 hours ago
Comment by inigyou 1 hour ago
Comment by nsbshsuzuh 8 hours ago
Comment by nathanappere 6 hours ago
Comment by DimitriBouriez 5 hours ago
Comment by TonyStr 5 hours ago
Comment by bulbar 3 minutes ago
Adding things that are not explicitly stated in the prompt but that are probable are the main benefit of AI, as far as I am concerned.
Comment by tjoff 4 hours ago
Comment by hn_throwaway_99 16 hours ago
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
Comment by viciousvoxel 15 hours ago
Comment by qwertybased 15 hours ago
edit: nevermind, definitely "monkey-puppet side-eye" vibe.
Comment by fennecfoxy 41 minutes ago
Comment by krisoft 13 hours ago
Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.
Comment by jostylr 12 hours ago
SVG: https://jostylr.com/imgs/frog_habsburg_jaw.svg
PNG: https://jostylr.com/imgs/frog_habsburg.png
Then I tried Codex with Sol 5.6 High and got a face forward one.
Comment by firesteelrain 10 hours ago
Comment by Rygian 4 hours ago
Comment by thebigship 16 hours ago
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
Comment by troupo 16 hours ago
That looks like something from Machinarium or Robots :)
Comment by abound 11 hours ago
Comment by uncivilized 10 hours ago
Comment by leptons 7 hours ago
Comment by riazrizvi 14 hours ago
Comment by evan_ 15 hours ago
Comment by 4sak3n 4 hours ago
Comment by wren6991 16 hours ago
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
Comment by vb-8448 4 hours ago
Comment by getnormality 16 hours ago
Comment by zirkuswurstikus 9 hours ago
ChatGPT MMD
Comment by ComputerGuru 8 hours ago
Comment by spencerflem 9 hours ago
Comment by rush86999 15 hours ago
That's a pretty good benchmark
Comment by ianberdin 15 hours ago
Comment by xenonite 5 hours ago
Comment by sn0n 4 hours ago
Comment by ricardobeat 15 hours ago
Comment by klooney 9 hours ago
Comment by dehrmann 16 hours ago
Comment by TSltd 12 hours ago
Comment by akomtu 11 hours ago
Comment by linksnapzz 16 hours ago
Comment by leumon 16 hours ago
Comment by thebigship 16 hours ago
Comment by leumon 15 hours ago
Comment by gerdesj 15 hours ago
Comment by yanhangyhy 6 hours ago
Comment by k1e 14 hours ago
Comment by thebigship 11 hours ago
Comment by MiroslavPokorny 15 hours ago
Comment by buffer_overlord 10 hours ago
Comment by konart 7 hours ago
But Kimi and Claude win this (from models listed on the page)
Comment by thebigship 6 hours ago
Comment by csomar 14 hours ago
I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.
Comment by throwuxiytayq 14 hours ago
Comment by cindyllm 14 hours ago
Comment by epolanski 14 hours ago
Comment by acdha 11 hours ago
That doesn’t mean there are no ways to use them productively but rather that you should keep in mind that the same model will happily give you code or a decision with the same level of error unless you have carefully setup a QA regimen to prevent that.
Comment by zahlman 5 hours ago
The standard retort seems to be "this is also true of a large proportion of humans". But I think it's clear that there are differing patterns in how humans vs. models err on various tasks.
Comment by acdha 2 hours ago
Comment by Mindless2112 10 hours ago
Comment by runarberg 14 hours ago
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.
Comment by viccis 14 hours ago
Comment by epolanski 15 hours ago
Would've wanted to see also DS4 flash.
Comment by troupo 16 hours ago
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
Comment by AlienRobot 14 hours ago
Comment by thebigship 17 hours ago
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.
Comment by zahlman 5 hours ago
You seem to imply that they ought not to. I disagree.
I wasn't familiar with the term before this post. Having learned it, were I given the task, I think I'd be strongly tempted to do the same extrapolation.
> If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.
Agency is agency. You still need to vet what the model's output is actually permitted to control.
Comment by thebigship 5 hours ago
That would be like saying anyone with Lou Gehrig's disease must look like Lou Gehrig. So we'll have to agree to disagree here.
Comment by n00bskoolbus 16 hours ago
Comment by fwip 15 hours ago
Comment by HPsquared 16 hours ago
Comment by thebigship 16 hours ago
Comment by sixtyj 16 hours ago
Comment by slipperybeluga 10 hours ago
Comment by kindawinda 15 hours ago
Comment by NemoNobody 14 hours ago