Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Posted by theanonymousone 6 hours ago
Comments
Comment by simonw 5 hours ago
I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.
I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.
Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Comment by Jimmc414 12 minutes ago
https://platform.claude.com/docs/en/models/opus-5-5/overview
Comment by zerof1l 3 hours ago
Comment by zozbot234 3 hours ago
Comment by dannyw 58 minutes ago
Comment by realusername 3 hours ago
I also switch to a better model for more complex tasks, also in low settings
Comment by RGS1811 5 hours ago
I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
Comment by croemer 4 hours ago
Comment by pgwhalen 1 hour ago
Comment by simonw 5 hours ago
Comment by dgellow 4 hours ago
Comment by 0x10ca1h0st 3 hours ago
Comment by cubefox 5 hours ago
Comment by Someone1234 5 hours ago
My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.
Comment by seabass-salmon 4 hours ago
Comment by arcanemachiner 2 hours ago
Have you ruled out the possibility that your system prompt, AGENTS.md, or increasing codebase complexity are not to blame?
Comment by samuelknight 5 hours ago
Comment by Gcam 1 hour ago
Comment by alansaber 3 hours ago
Comment by Izmaki 2 hours ago
Comment by sidewndr46 5 hours ago
I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.
Comment by az226 5 hours ago
Comment by simonw 5 hours ago
Piping the visible reasoning trace through their token counter API (I use https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.
Comment by beardsciences 5 hours ago
Comment by breckenedge 5 hours ago
Comment by tedsanders 31 minutes ago
GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.
We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.
(I work at OpenAI.)
Comment by breckenedge 7 minutes ago
Yesterday, I ran an identical bug identification dataset from two weeks ago, saw a 50% drop from a few weeks ago, putting Sol on the same level as Luna. Sol had been finding 40-50 bugs per set, then dropped to 25, matching Luna’s performance. Not enough to establish a pattern, but enough to raise eyebrows.
Our review workflow is public if you want to peruse it, dataset isn’t. The process isn’t really stabilized yet either as I have to balance running this against limited budgets.
https://github.com/BiggerPockets/.github/blob/main/.github/w...
Comment by tedsanders 1 minute ago
If it's a single task where it dropped from 50 to 25, it could be random variation. If it's the mean over hundreds of tasks, that suggests a problem with either the eval code/harness or our API.
Comment by mnicky 3 hours ago
Comment by echelon 3 hours ago
Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.
Comment by hglaser 5 hours ago
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
Comment by tomjakubowski 2 hours ago
Comment by sharktheone 5 hours ago
Comment by giancarlostoro 5 hours ago
Comment by asdfasgasdgasdg 3 hours ago
Comment by user43928 5 hours ago
Comment by onlyrealcuzzo 4 hours ago
It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.
I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.
An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.
Comment by makeavish 5 hours ago
Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level
Comment by linuxrebe1 3 hours ago
Comment by mckirk 18 minutes ago
Comment by mchusma 5 hours ago
Many benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.
Comment by ____tom____ 3 hours ago
That says something about your selected range, and nothing about the model.
Comment by TomGarden 30 minutes ago
Comment by sharktheone 5 hours ago
Comment by giancarlostoro 5 hours ago
Comment by khalic 2 hours ago
-_-‘
Comment by linuxrebe1 3 hours ago
Comment by qsort 5 hours ago
The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?
Comment by kzrdude 5 hours ago
Comment by losvedir 5 hours ago
So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.
Comment by ruszki 4 hours ago
Comment by Someone1234 5 hours ago
"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.
Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.
Comment by svachalek 2 hours ago
Comment by esafak 5 hours ago
Comment by qsort 5 hours ago
One man's modus ponens is another's modus tollens I guess.
Comment by bkishan 5 hours ago
Comment by meric_ 5 hours ago
(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)
Comment by firemelt 5 hours ago
can anyone help me?
Comment by floki165 4 hours ago
Comment by justindotdev 5 hours ago