Claude Opus 5.5
Posted by km144 6 hours ago
Comments
Comment by sailingparrot 6 hours ago
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
Comment by mukmuk 6 hours ago
Comment by DiggyJohnson 6 hours ago
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
Comment by mpalczewski 5 hours ago
Comment by mitchdoogle 2 hours ago
Comment by epolanski 1 hour ago
The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
Comment by usef- 45 minutes ago
I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
(I suspect not many people read the essay, judging by how many people seem surprised they're releasing improved models)
Comment by arrowleaf 17 minutes ago
Comment by glenstein 5 hours ago
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
Comment by cgio 2 hours ago
Comment by alwillis 2 hours ago
From https://en.wikipedia.org/wiki/Safety_car
> In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.
Comment by post-it 6 hours ago
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
Comment by sigmar 5 hours ago
Comment by antod 4 hours ago
eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.
Comment by smelendez 4 hours ago
Comment by bee_rider 3 hours ago
Comment by pegasus 3 hours ago
Comment by InsideOutSanta 2 hours ago
Comment by heroiccocoa 1 hour ago
Comment by 1attice 1 hour ago
It was still shit tier comms for communicating with the whole planet, but yes, for the inner loop, it was succinct and clear.
Valley neuralese
Comment by rcxdude 33 minutes ago
Comment by Avicebron 12 minutes ago
Comment by nradov 5 hours ago
Comment by TeMPOraL 4 hours ago
So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.
Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.
Comment by idiotsecant 4 hours ago
You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.
Comment by victorhooi 2 hours ago
1. It's just bad communication, full stop - just look at the comments here, even people allegedly in support of Anthropic are all arguing over what the phrase is even meant to mean.
2. It's flowery language and oddly out of place - which yes, can be triggering for people who have to deal with Claude doing this as well.
Claude seems overly apt to reach for "coinages", or neologism (yes, aha, I learnt that phrase, after spending time dealing with Claude...). It will create some made-up phrase to describe an otherwise dry, scientific CS concept, and nobody seems to know why. Surely it can't be user-focus groups?
So it would be peak-AI if somehow, the Anthropic communications team was also using Claude to author these blog posts, about how they were "pacing the frontier" - which either means they're betting big on AI, and going at it faster than OpenAI...or maybe it means they need to slow down releases, because it's too buggy...or maybe it means they're worried about regulatory capture? I honestly have no idea.
It's like the whole "Advancing Our Amazing Bet" corporate-speak from my old bosses - maybe they were trying to soften the blow or something, or be nice, but it ended up just confusing the heck out of everybody.. (Spoiler alert - the phrase actually meant they were shutting the whole thing down)
Comment by nradov 4 hours ago
Comment by freejazz 2 hours ago
Comment by johnisgood 5 hours ago
Is this the meaning or do I have it wrong? I have not checked.
Comment by lxgr 5 hours ago
Comment by ventana 3 hours ago
Comment by testdelacc1 2 hours ago
Comment by bityard 5 hours ago
Comment by wren6991 5 hours ago
Comment by johnisgood 5 hours ago
Comment by TeMPOraL 4 hours ago
Comment by johnisgood 2 hours ago
Comment by melasadra 5 hours ago
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
Comment by rhet0rica 5 hours ago
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
Comment by LanceH 5 hours ago
Comment by DiggyJohnson 3 hours ago
Comment by ck2 5 hours ago
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
Comment by tetha 5 hours ago
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
Comment by derac 5 hours ago
Comment by qlte 5 hours ago
Comment by stagger87 2 hours ago
No need to assume, the phrase is literally a link to the blog post the defines it!
Comment by irpap 1 hour ago
Comment by JackFr 2 hours ago
Comment by doctoboggan 4 hours ago
Comment by jgwil2 4 hours ago
Comment by adrianmonk 3 hours ago
The current situation with AI is that everyone is going as fast as possible. So, we can logically eliminate speeding up because it's impossible by definition. And we can practically eliminate staying the same speed because why make a big fanfare and coin a special term to announce that you're keeping the status quo. By process of elimination, it must mean slowing down.
Comment by cgriswald 46 minutes ago
Comment by arw0n 5 hours ago
Comment by victorhooi 2 hours ago
I know you said it sounds poetic...but your comment reinforced the parent's point - that this sort of flowery LLM-ish speech is just bad communication.
It would be equivalent of my taking say random quotes from, Romance of the Three Kingdoms, and trying to use it to explain to my boss why I didn't finish the TPS reports last night.
Or quoting Pablo Neruda, into a report about wheat futures pricing this week, and how it's like a voyage with waters and stars...(no I'm not going to quote the original Spanish, I'd simply mangle it).
(To be clear - this isn't a dig at you, as a non-native speaker - I'm simply pointing out that this sort of AI phrasing is often counterproductive).
Comment by hencq 5 hours ago
Comment by fragmede 5 hours ago
Comment by icedchai 1 hour ago
Comment by irpap 1 hour ago
Comment by nradov 5 hours ago
Comment by staindk 5 hours ago
Comment by squidbeak 5 hours ago
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
Comment by wavewrangler 1 hour ago
Comment by browningstreet 5 hours ago
Comment by logifail 5 hours ago
It's a strategy to achieve more, not less.
Comment by browningstreet 4 hours ago
A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.
Comment by lxgr 5 hours ago
Comment by neo_doom 5 hours ago
Comment by Leynos 5 hours ago
Comment by geraneum 45 minutes ago
Comment by patcon 5 hours ago
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
Comment by platinumrad 5 hours ago
Comment by DiggyJohnson 5 hours ago
Comment by lxgr 5 hours ago
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
Comment by chickensong 3 hours ago
Comment by switchbak 56 minutes ago
Why do you think it's your responsibility to police who says what about some megacorp, on a random forum on the internets? If I want to criticize some corporation's PR output, I think I'll go right ahead and do that, thanks.
Comment by jp57 3 hours ago
Now a computer scientist might claim that this use of "to pace <something>" is just a generalization of "to pace oneself", but as with many reflexive verb uses, there isn't really an equivalent usage with a non-reflexive object. It's kind of an invention. It's not necessarily wrong to invent a new usage, but usually one does it when there isn't really any other more direct way of saying it, and I don't think that's the case here.
Comment by aesthesia 1 hour ago
https://www.merriam-webster.com/dictionary/pace#dictionary-e... https://en.wiktionary.org/wiki/pace#Verb
Comment by BobbyJo 4 hours ago
There is almost always a large amount of time and effort invested behind the scenes in exactly how to message things like this. That being the case, there is almost always some insight to be had criticizing and analyzing what they settled on.
Comment by ben_w 3 hours ago
If I'm being cynical, "pacing" may sound nice, but "fast pace" and "slow pace" are both "pacing".
Comment by kelnos 5 hours ago
It's a weird phrase. Not sure why there are so many people who feel the need to defend it with such passion.
Comment by chickensong 3 hours ago
LLMs have made people so sensitive to language that I fear we're going to throw the baby out with the bath water. The models obviously need work, but they're also a great opportunity to expand our own vocabulary and grammar. It would be a shame if we deny some of the finer points of language in favor of Grug-speak to appease the lowest common denominator.
Comment by adrianmonk 2 hours ago
Comment by vmnb 5 hours ago
Comment by switchbak 1 hour ago
Comment by kadushka 5 hours ago
Comment by vasco 5 hours ago
Comment by fragmede 5 hours ago
Comment by alwillis 2 hours ago
Opus 5.5 is no closer to RSI than Opus 5 was.
Comment by janalsncm 4 hours ago
Comment by isoprophlex 5 hours ago
Comment by DiggyJohnson 5 hours ago
Comment by Citizen_Lame 3 hours ago
Comment by switchbak 1 hour ago
Seriously though, I can't believe people care this much about a stupid phrase - either for or against.
Comment by rythmshifter 5 hours ago
sir, this is a hacker news thread
Comment by swader999 1 hour ago
Comment by exe34 47 minutes ago
Comment by OJFord 5 hours ago
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
Comment by freejazz 2 hours ago
That's not clear at all. How could that be what "pacing" clearly means in the context of the "frontier". How is that more clear than any other pace that could be at issue???
Comment by lukewarm707 5 hours ago
"there is nothing outside the text" - Jacques Derrida
Comment by switchbak 59 minutes ago
Comment by kingkawn 3 hours ago
Comment by icedchai 4 hours ago
Comment by ltbarcly3 1 hour ago
If you think it clearly means anything you are just assuming because it can't "clearly" mean something specific when they go out of their way to use non idiomatic language and they don't give very clear guidance using idiomatic language.
Comment by Ar-Curunir 5 hours ago
Comment by DiggyJohnson 5 hours ago
Comment by tclancy 5 hours ago
Comment by usef- 7 minutes ago
The metaphor seems to be like a pacer runner in marathons: If you run too hard in the beginning of a marathon you will blow up and fail, so runners follow a pacer at the speed they can actually maintain safely.
Note that they wont necessarily be slower at finishing the overall race.
Comment by mpalczewski 5 hours ago
Comment by tclancy 3 hours ago
Comment by nonethewiser 5 hours ago
Comment by nradov 4 hours ago
https://www.war.gov/News/News-Stories/Article/Article/264106...
Comment by topbanana 5 hours ago
Comment by TheIronYuppie 4 hours ago
if you are in a long race, you don't run all out teh entire time. you pace yourself.
https://en.wikipedia.org/wiki/Pacing_strategies_in_track_and...
That couldn't be more exactly what they are doing here.
Comment by marton78 5 hours ago
Comment by Dumblydorr 5 hours ago
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
Comment by Rebelgecko 5 hours ago
Comment by lkbm 5 hours ago
"Pace yourself" specifically means "slow down".
Comment by nradov 3 hours ago
Comment by usef- 10 minutes ago
Comment by plaidfuji 5 hours ago
But stating it plainly like this would make the contradiction too obvious.
Comment by palmotea 5 hours ago
Claude says it sounds fine. And Claude is now the judge of the English language style, not you.
Comment by jugg1es 3 hours ago
Comment by jrochkind1 4 hours ago
It is clear what it means anyway, that's true, it means the left out words, more or less.
And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.
Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.
Comment by fluidcruft 4 hours ago
Comment by grohan 5 hours ago
Comment by gradus_ad 5 hours ago
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
Comment by apitman 3 hours ago
Comment by johnfn 47 minutes ago
Comment by usef- 32 minutes ago
Comment by SV_BubbleTime 1 hour ago
China is literally only a single step behind and willing to drop free models just to undercut the US companies.
I’m for it because I don’t want another massive Google or Meta.
Comment by dmazin 6 hours ago
Comment by sailingparrot 6 hours ago
Comment by davrosthedalek 6 hours ago
Comment by jr3592 6 hours ago
Comment by sailingparrot 6 hours ago
Comment by quietbritishjim 4 hours ago
Comment by jr3592 5 hours ago
Comment by sailingparrot 5 hours ago
Comment by lantry 5 hours ago
Comment by skerit 5 hours ago
Comment by recursive 5 hours ago
Comment by sidrag22 5 hours ago
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
Comment by sailingparrot 5 hours ago
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
Comment by sidrag22 5 hours ago
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
Comment by sailingparrot 1 hour ago
Comment by nicwolff 3 hours ago
Comment by dmix 6 hours ago
Comment by cab648bec139cc 6 hours ago
Comment by supern0va 6 hours ago
Comment by sleazebreeze 6 hours ago
Comment by re-thc 6 hours ago
Comment by cab648bec139cc 6 hours ago
Comment by anthonyrstevens 6 hours ago
Comment by felixgallo 5 hours ago
Comment by meowface 5 hours ago
Comment by 0xbadcafebee 6 hours ago
Comment by reasonableklout 5 hours ago
Comment by the_gipsy 6 hours ago
Comment by dgellow 6 hours ago
Comment by rubslopes 4 hours ago
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
Comment by jayd16 3 hours ago
Comment by usewik 5 hours ago
Comment by dgellow 4 hours ago
Comment by the_gipsy 4 hours ago
Comment by someothherguyy 6 hours ago
doesn't sound like a razor at all
Comment by bpodgursky 6 hours ago
Comment by nextaccountic 6 hours ago
Comment by ChrisLTD 4 hours ago
Comment by the_gipsy 4 hours ago
Comment by andkenneth 1 hour ago
Comment by tencentshill 5 hours ago
Comment by PaulStatezny 3 hours ago
I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.
Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.
So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:
1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.
2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."
Am I missing another likely option?
To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.
Comment by felixgallo 5 hours ago
Comment by staticman2 4 hours ago
Comment by felixgallo 1 hour ago
Comment by qgin 5 hours ago
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
Comment by 2001zhaozhao 11 minutes ago
For an idea of what a serious AI forecaster expects a coordinated AI slowdown to be feel like for the average citizen, see:
https://ai-2040.com/?choices=plan-a-root#playbook-public-pov
Comment by tantalor 5 hours ago
Comment by bonesss 4 hours ago
Comment by Tade0 4 hours ago
Comment by xadhominemx 2 hours ago
Comment by jatora 4 hours ago
Comment by xadhominemx 2 hours ago
Comment by Iolaum 6 hours ago
Comment by lukewarm707 6 hours ago
that, they fully intend to 'pace'.
Comment by drnick1 5 hours ago
Comment by SOLAR_FIELDS 3 hours ago
Comment by cgio 2 hours ago
Comment by lukewarm707 5 hours ago
what anthropic have stolen they intend to keep for themselves.
Comment by azan_ 6 hours ago
Comment by user3939382 6 hours ago
Comment by RcouF1uZ4gsC 14 minutes ago
Comment by kadushka 6 hours ago
Comment by msikora 1 hour ago
This term is quite ambiguous. Did Dario mean that they need to go faster while making it sounds like they will slow down???
Comment by Lendal 4 hours ago
Comment by hnha 4 hours ago
If their scare was honest, they would stop.
Comment by AgentME 3 hours ago
Comment by jdale27 2 hours ago
Comment by bigfishrunning 2 hours ago
therefore, their scare is marketing.
Comment by maxutility 3 hours ago
Comment by rudedogg 3 hours ago
And I don’t think any pacing is/was intentional. They’de release skynet if they could and the stonks went up
Comment by sailingparrot 3 hours ago
Comment by mullingitover 5 hours ago
Comment by DonsDiscountGas 2 hours ago
Comment by scottyah 6 hours ago
Comment by wavewrangler 1 hour ago
Comment by tombert 2 hours ago
"Our technology is so unbelievably powerful that the entire world might shatter if we don't have government imposed handcuffs!!!!". It just reads like the corporate equivalent of the drunk frat guy saying "HOLD ME BACK BRO!"
Comment by dspillett 5 hours ago
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
Comment by guybedo 2 hours ago
Opus 5.5 isn't the frontier, when they say 'pacing the frontier', it's about internal models not yet released, as they're probably one or two generations ahead already.
Comment by chinathrow 4 hours ago
Comment by dr0idattack 6 hours ago
Comment by janpot 5 hours ago
Comment by heyjstn 5 hours ago
Comment by AtlasBarfed 5 hours ago
Simply make them something that derives a text response from its training data.
Comment by BatmansMom 6 hours ago
Comment by sailingparrot 6 hours ago
Comment by whalesalad 4 hours ago
Comment by CodingJeebus 6 hours ago
Comment by dboreham 20 minutes ago
Comment by jr3592 5 hours ago
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
Comment by varispeed 4 hours ago
Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.
Idiotic.
Comment by Flere-Imsaho 2 hours ago
https://openai.com/index/introducing-gpt-6-sol-and-luna/
Yeah think I'll be using OpenAI/Deepseek/etc from now on. I don't need your model to decide for me what is and isn't safe.
Comment by GodelNumbering 6 hours ago
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
Comment by AJ007 6 hours ago
Comment by gwd 1 hour ago
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
Comment by retinaros 5 minutes ago
Comment by mcintyre1994 5 hours ago
Comment by drbscl 5 hours ago
It does work out to be a similar cost per task though
Comment by jsnell 5 hours ago
https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
Comment by drbscl 5 hours ago
Comment by persedes 2 hours ago
Comment by 93po 4 hours ago
Comment by piotrdz 3 hours ago
Comment by johnbellone 3 hours ago
Comment by epolanski 1 hour ago
Comment by naasking 5 hours ago
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
Comment by meerita 50 minutes ago
Comment by make3 5 hours ago
Comment by _the_inflator 5 hours ago
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
Comment by blfr 6 hours ago
Comment by rapfaria 6 hours ago
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
Comment by blfr 6 hours ago
Comment by herpdyderp 5 hours ago
Comment by girvo 1 hour ago
Comment by neuronexmachina 5 hours ago
Comment by ascorbic 5 hours ago
Comment by coffeebeqn 6 hours ago
Comment by btown 5 hours ago
Comment by chrisweekly 5 hours ago
Comment by cute_boi 5 hours ago
Comment by chrisweekly 4 hours ago
In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.
Comment by notatoad 5 hours ago
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
Comment by margorczynski 4 hours ago
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
Comment by johnecheck 4 hours ago
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
Comment by bulbar 3 hours ago
They will burn as much money as necessary to make that happen. And they have a virtually infinite amount of liquidity.
Comment by epolanski 1 hour ago
There's only 12 countries that do on the planet, the most "relevant" of them being Guatemala and Haiti.
As or the LLM topic: you can download weights of chinese models and remove any censorship and bias. Can you do so with american closed ones?
Comment by dboreham 19 minutes ago
Comment by Shekelphile 5 hours ago
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
Comment by andxor 3 minutes ago
Comment by brookst 5 hours ago
Comment by artursapek 1 hour ago
Comment by alvis 6 hours ago
Comment by bayesianbot 6 hours ago
Comment by weiran 6 hours ago
Comment by Espressosaurus 6 hours ago
Comment by hedgehog 5 hours ago
Comment by re-thc 6 hours ago
For long running tasks it is. That's what made Deepseek so cheap.
Comment by vardalab 5 hours ago
Comment by liudaisuda 6 hours ago
Comment by forgot-my-pw 4 hours ago
Comment by rahimnathwani 6 hours ago
Comment by mcintyre1994 5 hours ago
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
Comment by derangedHorse 5 hours ago
Comment by BatFastard 5 hours ago
Comment by comboy 4 hours ago
Comment by atonse 2 hours ago
Not a day goes by when I push back on something, to which Opus 5 very unambiguously say "You were right, I was wrong" - this never happened so often with past models, nor with Fable.
We'll have to see how much Opus's ability to communicate has improved. It's already giving me better summaries of where we are in the conversation.
Comment by jaflo 5 hours ago
Comment by itsafarqueue 5 hours ago
Comment by mitchdoogle 2 hours ago
Comment by mcintyre1994 4 hours ago
Comment by weego 1 hour ago
Navigating the landscape of agentic levers certainly requires a more detailed approach than this and you were certainly correct to push back.
Comment by mcintyre1994 42 minutes ago
Comment by UnboundedContex 1 hour ago
Comment by physicles 2 hours ago
Fable 5.1 is a lot better than Fable 5 btw (edit: in terms of writing style). Not sure about opus 5.5 yet since I’ve only got one session in so far.
Comment by epicepicurean 5 hours ago
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
Comment by croemer 4 hours ago
Comment by throwaway219450 3 hours ago
> There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
Comment by pgphn 3 hours ago
Comment by californical 4 hours ago
> Rewrites are declared by the publisher, never inferred from overlap
> NULL means dirty, and DELETE is the fence
Comment by croemer 4 hours ago
Comment by rfgplk 3 hours ago
> Rewrites are declared by the publisher, never inferred from overlap.
This style of writing is idiotic because it conveys no additional information. It's no different from stating
> Rewrites are declared by the publisher, never when moons collide.
The two sentences are actually logically identical. No idea why these models keep writing like this.
> Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier.
This is even more ridiculous.
Comment by altern8 3 hours ago
Comment by algoth1 5 hours ago
Comment by sha-3 5 hours ago
Comment by Trasmatta 5 hours ago
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
Comment by nonethewiser 5 hours ago
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
Comment by penagwin 4 hours ago
That’s the step that causes the most significant gains in agentic performance.
But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).
That’s why it often gets worse on models that simply had more RL post training from the same base.
Comment by nonethewiser 3 hours ago
Comment by tancop 58 minutes ago
Apparently it helps generalize skills between areas, which makes sense when you compare it to how humans learn but I don't know if it's the same for LLMs.
Comment by LtdJorge 5 hours ago
Comment by rfgplk 3 hours ago
Comment by Aperocky 5 hours ago
Comment by FireBeyond 4 hours ago
Comment by legobmw99 2 hours ago
Comment by epolanski 45 minutes ago
People that produce slop have to be fired asap, they're just human relays anyway.
Comment by wg0 3 hours ago
I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.
Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.
My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.
Total cost of the above? $0.07 cents.
PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.
Comment by glub 2 hours ago
Comment by wg0 1 hour ago
Comment by tontinton 2 hours ago
Comment by wg0 2 hours ago
And that all is 0.07 cents all included.
Comment by s3p 2 hours ago
Comment by wg0 2 hours ago
PS: I do not know why but opencode pushes CPU usage to very high which has NOT happened with DeepSeek harness even once.
Comment by meerita 1 hour ago
Comment by simonw 5 hours ago
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.mdComment by MikhailTal 5 hours ago
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
Comment by Brendinooo 5 hours ago
Comment by nonethewiser 5 hours ago
"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."
Comment by copperx 5 hours ago
Comment by MaxikCZ 5 hours ago
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Comment by simonw 5 hours ago
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Comment by zamadatix 5 hours ago
Comment by segbrk 5 hours ago
Comment by FergusArgyll 5 hours ago
Comment by nijave 5 hours ago
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
Comment by ceroxylon 4 hours ago
Comment by adverbly 4 hours ago
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
Comment by DenisM 2 hours ago
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
Comment by mjhagen 44 minutes ago
Comment by ilaksh 2 hours ago
Comment by TomGarden 3 hours ago
I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.
Comment by Kailhus 1 hour ago
Comment by ealready_value 5 hours ago
Comment by cainxinth 5 hours ago
Comment by Kurtz79 5 hours ago
Comment by caxco93 2 hours ago
Comment by inshard 5 hours ago
Comment by skerit 5 hours ago
Comment by spidersouris 4 hours ago
Comment by breezybottom 4 hours ago
Comment by nicolamanzini 3 hours ago
Comment by ipsum2 2 hours ago
Comment by nicolamanzini 1 hour ago
Comment by PetahNZ 3 hours ago
Comment by make3 5 hours ago
Comment by hamrocksissors 3 hours ago
Comment by copperx 5 hours ago
Comment by ApolloFortyNine 6 hours ago
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
Comment by sys32768 5 hours ago
ChatGPT 6 Pro answered it without issue.
Comment by debesyla 4 hours ago
Comment by timacles 4 hours ago
Comment by dopa42365 3 hours ago
Comment by b112 3 hours ago
Comment by toss1 4 hours ago
So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.
[0] https://www.sciencebuddies.org/projects-lessons-activities/g...
[1] https://www.genewiz.com/public/services/sanger-sequencing
Comment by peri-cl 5 hours ago
https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...
Comment by blfr 6 hours ago
Comment by arw0n 5 hours ago
Comment by kqp 4 hours ago
Comment by cute_boi 5 hours ago
Giving moral lecture is different than reality i guess.
Comment by prettyblocks 6 hours ago
Comment by Espressosaurus 6 hours ago
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
Comment by raesene9 5 hours ago
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
Comment by flyinglizard 4 hours ago
Comment by Metacelsus 5 hours ago
Comment by doginasuit 3 hours ago
Comment by TuxSH 14 minutes ago
Notice that this isn't cybersec nor memory-safety related at all.
Comment by nonethewiser 5 hours ago
I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.
Comment by paimapi 2 hours ago
Comment by bushido 6 hours ago
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
Comment by ACCount39 5 hours ago
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
Comment by user43928 2 hours ago
>Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.
But hey, they 'should not impact the vast majority' of ML development. Great.
Comment by dannyw 46 minutes ago
Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.
Comment by KeplerBoy 5 hours ago
Comment by yaakov34 3 hours ago
Comment by SoftTalker 4 hours ago
Comment by searine 6 hours ago
Comment by unglaublich 6 hours ago
Comment by nijave 5 hours ago
Comment by b112 3 hours ago
So there are literal avenues to identify yourself, very cheaply, with a human. Theoretically, a company with its own AI, should be able to support more than just Persona, after all.. SDK integration should be simplistic for them.
Anthropic? Support domestic eID providers, you can even use it as advertising "See how easy AI makes it?" and "We care!" and so forth.
At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.
Comment by techjamie 6 hours ago
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
Comment by stri8ted 5 hours ago
Comment by manquer 4 hours ago
Comment by dannyw 45 minutes ago
Comment by Balinares 3 hours ago
Comment by ACCount39 5 hours ago
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
Comment by ryangg 6 hours ago
Comment by peri-cl 5 hours ago
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
Comment by potwinkle 5 hours ago
Comment by zatkin 5 hours ago
Comment by joshstrange 6 hours ago
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
Comment by bleonard 5 hours ago
So longer threads get cheaper and one-shots stay the same price.
Comment by bayesianbot 6 hours ago
Comment by m4tthumphrey 6 hours ago
Comment by amluto 6 hours ago
Comment by gruez 6 hours ago
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
Comment by KyleTheDev 6 hours ago
I agree that it's sort of stupid, not a fan.
Comment by ealready_value 5 hours ago
Comment by EricBurnett 6 hours ago
Comment by mbreese 6 hours ago
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
Comment by thejazzman 6 hours ago
Comment by giancarlostoro 6 hours ago
Comment by iAMkenough 6 hours ago
Everyone that doesn't gets served some animated bullshit.
Comment by gruez 5 hours ago
Yep, you're right. I tried on my phone and got the scroll through image.
Comment by swader999 6 hours ago
Comment by serchinastico 5 hours ago
Comment by thebitguru 6 hours ago
Comment by halyconWays 6 hours ago
Comment by josefresco 6 hours ago
Comment by oefrha 5 hours ago
Comment by halyconWays 5 hours ago
Comment by iAMkenough 6 hours ago
Comment by dionian 6 hours ago
Comment by somewhatjustin 6 hours ago
Nice. I was starting to think that Haiku got abandoned.
Comment by Sol- 6 hours ago
Comment by ricardobeat 5 hours ago
Comment by w-m 1 hour ago
When I had Sol orchestrate Luna and Terra as implementation agents, Sol was a lot happier with what Terra produced and would find far fewer issues than what was implemented by Luna.
But a few weeks after introduction, OpenAI slashed Luna's cost by 80% and Terra's only by 20%. Only then did it become uneconomical to run Terra and its reason to exist stopped.
Comment by skerit 4 hours ago
Comment by bix6 3 hours ago
Comment by somewhatjustin 6 hours ago
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
Comment by jaapz 3 hours ago
Comment by mchusma 6 hours ago
Comment by cesarvarela 6 hours ago
Comment by sharkjacobs 6 hours ago
God I hope so
Comment by lgessler 6 hours ago
Comment by nonethewiser 4 hours ago
Comment by fastball 6 hours ago
Comment by drbscl 5 hours ago
Comment by Trasmatta 5 hours ago
Comment by mikeocool 6 hours ago
Comment by kantahayashi 6 hours ago
Comment by bushido 6 hours ago
Comment by bkishan 5 hours ago
Comment by rfgplk 3 hours ago
Comment by boc 6 hours ago
Comment by unddoch 5 hours ago
Comment by neilellis 5 hours ago
Comment by mavamaarten 6 hours ago
Comment by abtinf 6 hours ago
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
Comment by cbg0 6 hours ago
Comment by qlte 5 hours ago
https://artificialanalysis.ai/models/releases/claude-opus-5-...
Opus 5.5 Medium = $1.34
GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
Comment by TuxSH 46 seconds ago
Comment by margorczynski 4 hours ago
Comment by abtinf 3 hours ago
The Claude lock-in simply disqualifies anthropic entirely (for my use).
Comment by copperx 5 hours ago
That's an incredibly bold assumption.
Comment by cbg0 5 hours ago
Comment by notatoad 4 hours ago
Comment by onlyrealcuzzo 6 hours ago
This is news to me. Excited to try it out! Thanks.
Comment by nchmy 6 hours ago
Comment by KeplerBoy 5 hours ago
Comment by polalavik 6 hours ago
Comment by abtinf 3 hours ago
Comment by roughly 6 hours ago
Can you give more details here? This sounds intriguing.
Comment by sidrag22 5 hours ago
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
Comment by mlcruz 5 hours ago
Comment by felixgallo 6 hours ago
Comment by abtinf 6 hours ago
Comment by ryanscio 6 hours ago
Comment by felixgallo 6 hours ago
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
Comment by esafak 5 hours ago
Comment by zuInnp 6 hours ago
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
Comment by orangecat 5 hours ago
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
Comment by ieie3366 6 hours ago
Comment by ACCount39 5 hours ago
Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.
Comment by glub 6 hours ago
Benchmarks often don't survive contact with reality.
Comment by drnick1 5 hours ago
Comment by cowthulhu 5 hours ago
Comment by cheikhcheikh 5 hours ago
Comment by drnick1 5 hours ago
Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.
Comment by arw0n 5 hours ago
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
Comment by boredtofears 5 hours ago
Comment by notatoad 5 hours ago
Comment by giancarlostoro 5 hours ago
Comment by Imustaskforhelp 5 hours ago
Comment by quotemstr 6 hours ago
Comment by sznio 6 hours ago
Comment by booty 6 hours ago
Comment by NorwegianDude 4 hours ago
The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.
Comment by copperx 5 hours ago
Comment by copperx 5 hours ago
Comment by lanyard-textile 6 hours ago
Comment by ygouzerh 6 hours ago
Comment by Zambyte 5 hours ago
Comment by mavamaarten 6 hours ago
Comment by adastra22 4 hours ago
Comment by system2 6 hours ago
Comment by enraged_camel 5 hours ago
Comment by anthonypasq 5 hours ago
Comment by enraged_camel 5 hours ago
Comment by Game_Ender 4 hours ago
Comment by mintik 1 hour ago
Comment by nullbio 21 minutes ago
Comment by throwaway2027 6 hours ago
Comment by lgessler 5 hours ago
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
Comment by jaapz 3 hours ago
Comment by handfuloflight 6 hours ago
Comment by rich_sasha 6 hours ago
Comment by cmrdporcupine 6 hours ago
Comment by danw1979 6 hours ago
Comment by staticman2 6 hours ago
Comment by danw1979 6 hours ago
Comment by ThouYS 6 hours ago
Comment by hmokiguess 6 hours ago
Comment by aoeusnth1 6 hours ago
Comment by Retr0id 3 hours ago
Comment by loopmonster 5 hours ago
Comment by sailfast 6 hours ago
Comment by cronin101 6 hours ago
Comment by carlos-menezes 6 hours ago
Comment by RGS1811 6 hours ago
Comment by bibimsz 5 hours ago
Comment by fghorow 6 hours ago
Comment by esafak 6 hours ago
Comment by tda 5 hours ago
Comment by shibel 1 hour ago
Comment by throwuxiytayq 1 hour ago
Comment by kibae 4 hours ago
This is where Chinese models are going to eat Anthropic's lunch.
Comment by mlh496 1 hour ago
So the lack of guardrails is a very risky proposition...
Comment by nomel 3 hours ago
Comment by the_doctah 56 minutes ago
Comment by AlfeG 4 hours ago
Comment by TomGarden 3 hours ago
Quoted:
"Please explain the issue to me.
Claude Opus 5.5:
The extra drop is a bug in the billing refactor
The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.
What changed
Before the merge, aggregate.py used a half-open interval: /.../ last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all."
Comment by rfgplk 3 hours ago
Comment by magicalhippo 3 hours ago
I don't have time to really get to know one model before the next is out, and I'm just talking about OpenAI and Anthropic, never mind the long tail of alternatives.
So I just more or less haphazardly pick one based on the mood I'm in, and set reasoning effort based on how much quota I have left.
Comment by 2001zhaozhao 5 hours ago
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
Comment by 6thbit 28 minutes ago
Comment by nullbio 25 minutes ago
Comment by mgw 6 hours ago
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
Comment by mudkipdev 6 hours ago
Comment by simianwords 6 hours ago
Comment by enraged_camel 6 hours ago
Comment by slowin 4 hours ago
Comment by sebastiangrill 4 hours ago
Comment by slowin 4 hours ago
Comment by garo-pro 6 hours ago
Comment by aesthesia 4 hours ago
Comment by johnmlussier 1 hour ago
Comment by senko 2 hours ago
Minecraft clone: https://senko.net/vibecode-bench/2026/voxel-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/voxel-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/voxel-gpt-6-astra.html (Astra 6)
Warcraft clone: https://senko.net/vibecode-bench/2026/rts-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/rts-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html (Astra 6)
The above Opus games took ~45min to generate with the cost between $11 and $14 (per ccusage - I'm on a Max sub). Used from Claude Code with xhigh effort.
Full tests with prompts: https://senko.net/vibecode-bench/
Comment by kitbrennan 1 hour ago
Comment by skunkworker 6 hours ago
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
Comment by ekckekcjekfj 6 hours ago
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
Comment by Gander5739 6 hours ago
Comment by tomhow 6 hours ago
Comment by km144 6 hours ago
Comment by tomhow 6 hours ago
Comment by tomaskafka 5 hours ago
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
Comment by jjcm 5 hours ago
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
Comment by copperx 5 hours ago
Comment by jjcm 4 hours ago
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
Comment by naet 5 hours ago
Comment by jjcm 4 hours ago
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
Comment by meerita 6 hours ago
Comment by meerita 53 minutes ago
Comment by aragornii 6 hours ago
Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.
Comment by ironqcold 1 hour ago
Comment by artursapek 59 minutes ago
Comment by m101 6 hours ago
Comment by nyx 6 hours ago
Comment by m101 5 hours ago
Comment by pavlov 6 hours ago
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
Comment by rumblefrog 6 hours ago
Comment by glub 6 hours ago
Comment by edude03 5 hours ago
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
Comment by Retro_Dev 6 hours ago
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
Comment by wren6991 5 hours ago
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
Comment by ACCount39 5 hours ago
An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.
This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.
The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.
Comment by wren6991 4 hours ago
Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?
Comment by ACCount39 4 hours ago
Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.
A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.
Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.
It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.
Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.
But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere. They wouldn't be nearly as advanced if they couldn't do it.
Comment by wren6991 3 hours ago
Comment by Retro_Dev 4 hours ago
Comment by villish 5 hours ago
That's the moat. Mistral has the capability but not the legal protections.
Comment by staticman2 2 hours ago
Couldn't I simply give a Chinese friend my key on Open router?
Comment by foltik 2 hours ago
I say just let them duke it out. After a decade of regulatory capture and enshittification, it’s nice to see some actual competition again.
Comment by ayhanfuat 6 hours ago
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
Comment by hugodan 48 minutes ago
Comment by sunaookami 2 hours ago
Comment by pookieinc 6 hours ago
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
Comment by jbellis 6 hours ago
Comment by meric_ 6 hours ago
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
Comment by randomblock1 6 hours ago
Comment by mosselman 4 hours ago
Comment by dezmou 1 hour ago
Comment by buntp 6 hours ago
Comment by frshgts 5 hours ago
Comment by wren6991 6 hours ago
Comment by FergusArgyll 5 hours ago
Comment by arendtio 2 hours ago
Why do those labs keep releasing on the same day?!?
Comment by herpdyderp 2 hours ago
Comment by throw03172019 2 hours ago
Comment by ieie3366 4 hours ago
Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with
Comment by Catloafdev 6 hours ago
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
Comment by dgroshev 6 hours ago
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
Comment by sashank_1509 6 hours ago
Comment by dgroshev 6 hours ago
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
Comment by wren6991 5 hours ago
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
Comment by dgroshev 5 hours ago
Claude is just comically bad nowadays.
Comment by redox99 5 hours ago
Comment by cruffle_duffle 3 hours ago
Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.
Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.
Comment by gekoxyz 6 hours ago
Comment by ithkuil 6 hours ago
Comment by mavamaarten 6 hours ago
Comment by aray07 6 hours ago
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
Comment by username_my1 6 hours ago
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
Comment by meric_ 6 hours ago
https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
Comment by ygouzerh 6 hours ago
Comment by Eliezer 5 hours ago
Comment by adastra22 4 hours ago
Comment by j_heffe 5 hours ago
Comment by brandon272 6 hours ago
Comment by gwking 6 hours ago
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
Comment by brandon272 6 hours ago
Comment by fragmede 4 hours ago
Comment by qurren 6 hours ago
Comment by drnick1 6 hours ago
Comment by ianberdin 2 hours ago
https://playcode.io/blog/macbook-svg-benchmark#model-claude-...
Btw, we have added Opus 5.5 as default model to playcode.ai
Comment by cogythea 6 hours ago
Comment by jdmoreira 6 hours ago
Comment by NielsHarksen 3 hours ago
Comment by bredren 6 hours ago
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
Comment by ryanscio 6 hours ago
Comment by benjiro29 6 hours ago
Comment by dbbk 6 hours ago
Comment by re-thc 6 hours ago
Comment by nozzlegear 6 hours ago
Comment by re-thc 4 hours ago
Comment by nozzlegear 59 minutes ago
Comment by petesergeant 6 hours ago
Comment by variety8675 6 hours ago
Comment by emadabdulrahim 6 hours ago
Comment by akhilome 6 hours ago
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
Comment by gavinray 6 hours ago
Comment by alpineman 6 hours ago
Comment by voiceeh 20 minutes ago
Comment by tomaskafka 3 hours ago
About the time.
Comment by ramoz 6 hours ago
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
Comment by bitexploder 5 hours ago
Comment by jatins 6 hours ago
Thank you.
Comment by calibas 6 hours ago
We can't test it properly because it knows it's being tested.
Comment by johntb86 6 hours ago
Comment by jwpapi 1 hour ago
I don’t care to look up terms as long as they are correct.
Comment by notduckrabbit 6 hours ago
Comment by manmal 6 hours ago
Comment by somewhatjustin 6 hours ago
Nice. I was starting to think Haiku was going to be abandoned.
Comment by glub 6 hours ago
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
Comment by madjam002 4 hours ago
It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.
It's frustrating that there isn't more transparency here.
Comment by blfr 5 hours ago
Comment by jacobgold 6 hours ago
Comment by aurareturn 6 hours ago
I tried Opus 5 and Astra.
Comment by toephu2 4 hours ago
Are the frontier labs even working on this problem?
Comment by mnicky 2 hours ago
Comment by toephu2 1 hour ago
even at xhigh I get context compaction quite a bit.
Comment by dom96 5 hours ago
It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.
Comment by lousken 5 hours ago
Comment by 34679 5 hours ago
Maybe this model can finally figure it out for them.
Comment by doodlesdev 5 hours ago
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
Big, if true.Comment by jdthedisciple 6 hours ago
Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?
How would this alleged difference (most likely bs) actually show up in reality?
GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.
Comment by enraged_camel 5 hours ago
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
Comment by jdthedisciple 3 hours ago
> how would this alleged difference (most likely bs) actually show up in reality?
Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright
All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?
Comment by Foobar8568 6 hours ago
Comment by breezybottom 5 hours ago
Not efficiency in writing, clearly.
Comment by thibran 6 hours ago
Comment by isodev 5 hours ago
Comment by spiderice 5 hours ago
Comment by isodev 5 hours ago
Comment by b38484848 5 hours ago
Comment by KasianFranks 2 hours ago
Comment by Retr0id 5 hours ago
Yay, yet another model I can't use for anything interesting, even with CVP.
Comment by desmondl 5 hours ago
Comment by km144 6 hours ago
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
Comment by booty 6 hours ago
"benchmark margins have become a less
reliable guide to real-world differences"
sounds like a big problem.
My guesses:1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
Comment by simianwords 6 hours ago
Comment by suddenlybananas 6 hours ago
Comment by CPLX 6 hours ago
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
Comment by port3000 6 hours ago
Comment by booty 6 hours ago
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
Comment by cbg0 6 hours ago
Comment by Syntaf 6 hours ago
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
Comment by jidaigeist 6 hours ago
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
Comment by andriy_koval 5 hours ago
Comment by b38484848 5 hours ago
Comment by the_gipsy 6 hours ago
Comment by datadrivenangel 6 hours ago
Comment by nimonian 3 hours ago
Comment by cruffle_duffle 3 hours ago
Comment by alvis 6 hours ago
Comment by greenavocado 6 hours ago
Comment by nickandbro 6 hours ago
Comment by keeganpoppen 6 hours ago
Comment by hi_hi 1 hour ago
Infuriating.
Comment by selcuka 14 minutes ago
Comment by __vivek 5 hours ago
Comment by yipinwong 5 hours ago
Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
Comment by copperx 5 hours ago
Comment by yipinwong 5 hours ago
---
I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.
I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).
---
As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.
Also adding verification for that Fable 5.1 output in the same sesssion.
Comment by toasty228 1 hour ago
Comment by tag2103 6 hours ago
Comment by HarHarVeryFunny 4 hours ago
Ants: It's a good model, sir!
Comment by iamsyr 6 hours ago
Comment by blurbleblurble 5 hours ago
Comment by aennassiri 6 hours ago
Comment by thatxliner 4 hours ago
Comment by keeeba 6 hours ago
Comment by velcrovan 6 hours ago
Comment by adastra22 4 hours ago
Comment by hadlock 5 hours ago
Comment by rs_rs_rs_rs_rs 4 hours ago
Comment by garo-pro 5 hours ago
Comment by richardjennings 6 hours ago
Comment by mococa 6 hours ago
Comment by kar1181 4 hours ago
Comment by woeirua 5 hours ago
Comment by Fizzadar 4 hours ago
Comment by LoganDark 6 hours ago
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
Comment by thinkingtoilet 3 hours ago
Comment by Yabood 6 hours ago
Comment by sandos 3 hours ago
Comment by Lord_Zero 6 hours ago
Comment by kingstnap 6 hours ago
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
Comment by firemelt 5 hours ago
Comment by cmrdporcupine 5 hours ago
https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...
Comment by karp773 5 hours ago
Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.
Comment by nimonian 3 hours ago
Comment by bdangubic 5 hours ago
Comment by ramesh31 5 hours ago
Comment by nailer 5 hours ago
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
Comment by anentropic 6 hours ago
Comment by Arcuru 6 hours ago
Comment by simianwords 6 hours ago
Comment by jdw64 6 hours ago
Comment by ricardobeat 6 hours ago
Comment by viccis 6 hours ago
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
Comment by scrollop 6 hours ago
Comment by phendrenad2 6 hours ago
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
Comment by snvzz 5 hours ago
Yup. As unusable as Fable 5.1, for assembly on 80s 68k personal computer platform. Awful.
Comment by blurbleblurble 5 hours ago
Comment by theGeatZhopa 5 hours ago
Comment by mupuff1234 6 hours ago
Comment by setsewerd 6 hours ago
Comment by icrbow 6 hours ago
Comment by petesergeant 6 hours ago
Comment by mupuff1234 6 hours ago
Less companies involved means less pressure to go fast.
Comment by nozzlegear 6 hours ago
Comment by roughly 6 hours ago
Comment by WarmWash 6 hours ago
Comment by Lord_Zero 6 hours ago
Comment by Madmallard 4 hours ago
chinese models can't come soon enough
we're already getting enshittification
Comment by vividfrier 5 hours ago
Comment by giancarlostoro 6 hours ago
Comment by xenit_v0 6 hours ago
Comment by Gander5739 6 hours ago
Comment by Helldez 3 hours ago
Comment by ace2pace 6 hours ago
Comment by nicolamanzini 4 hours ago
Comment by mrbonner 5 hours ago
Comment by SadErn 6 hours ago
Comment by gopalv 6 hours ago
Comment by hirako2000 6 hours ago
Infomercial at its best.
No wonder we are hammered with ai announcements.