Qwen3.8 Max now ranked as the best overall model by agentic index
Posted by apitman 10 hours ago
Comments
Comment by jjcm 8 hours ago
What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
Comment by monster_truck 6 hours ago
When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left
And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome
Comment by 4d4m 2 hours ago
Comment by aliasxneo 6 hours ago
I suspect that the models are genuinely close and that certain experiences get felt across providers but are inconsistent enough to convince people one is superior to the other. I for one have tried Deepseek on and off since my co-founder is fond of it and I've stopped trying now because I never have a good experience.
Comment by wanderlust123 5 hours ago
Comment by kadushka 5 hours ago
Comment by Cookingboy 5 hours ago
There would be no reason to if you are in the privileged position where cost isn't an issue.
For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.
Comment by FuckButtons 3 hours ago
Comment by kadushka 2 hours ago
Comment by esafak 4 hours ago
Comment by matheusmoreira 2 hours ago
Self-hosting is the biggest reason.
Comment by Systemerror7A69 1 hour ago
Comment by cyanydeez 5 hours ago
Comment by 14u2c 2 hours ago
Comment by pessimizer 5 hours ago
Comment by TacticalCoder 4 hours ago
Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians.
We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war.
Or we could talk about the number of nice palestinians killed since the beginning of the war in Gaza. Or we could go a bit further and talk about the joy and celebration in Gaza after their heroes brought back 200 hostages after having slaughtered 1200 civilians.
You may be living in a place that you think shields you from those but I know the ideologies behind these acts.
The fallacy of gray is just that: it's not true that there's always a nice middle ground and that there's no evil ideology out there.
Something something about the price of liberty being eternal vigilance. For there are people abusing your blind trust.
Comment by tovlier 3 hours ago
Comment by peterashford 3 hours ago
Comment by swat535 4 hours ago
I think it's mainly due to poor education many receive and a very controlled media that suppresses information.
It's shocking considering how much money they spend on education compared to other nations.
Comment by nimchimpsky 2 hours ago
Comment by sieabahlpark 5 hours ago
Comment by doginasuit 7 hours ago
Comment by AustinDev 7 hours ago
Comment by conception 4 hours ago
Comment by FuckButtons 3 hours ago
Comment by ofjcihen 6 hours ago
Comment by michelsedgh 7 hours ago
Comment by AussieWog93 7 hours ago
Comment by dullcrisp 3 hours ago
Comment by ygjb 1 hour ago
This isn't an anti-American sentiment. It is an anti-corporate/regulatory capture/embrace and extinguish sentiment (which probably reads the same to many people these days).
Comment by Gigachad 7 hours ago
Comment by BeetleB 6 hours ago
But they didn't find it. The Big LLM provider accepted guilt and paid a fine.
You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
Comment by kennywinker 6 hours ago
Comment by TheOtherHobbes 4 hours ago
Training from copies has been ruled fair use because it's "transformative" and not simply "derivative."
This is obviously debatable, but that's where the debate is at the moment.
Comment by kennywinker 1 hour ago
Because of the rulings of a couple of judges. Is that actually what the majority of people think?
> Copyright law only considers illegal ownership of a work
That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing.
Similarly, copyright has something to say if I read a legal copy of harry potter and then create a new work in that world.
Comment by michelsedgh 4 hours ago
Comment by BeetleB 2 hours ago
Because that use case is actually permitted by law.
Comment by kennywinker 1 hour ago
The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage.
So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.
Comment by nolok 6 hours ago
That's not how it works. You have to give it back.
Otherwise, the distiller can just pay a fine (no larger than the original did) and be okay then, right ?
Comment by BeetleB 2 hours ago
Comment by throwaway27448 5 hours ago
Comment by mannanj 6 hours ago
Comment by CMay 6 hours ago
So instead of relying heavily on human bottlenecks, you focus on agentic task verification since that's the low hanging fruit and verifiable at scale?
Comment by edg5000 1 hour ago
LLMs are certainly more knowledgable, but maybe not more intelligent, arguably. It's possible we're approacing a ceiling indeed.
Model capability might be on an asymptote appraching but never quite reaching parity with human intelligence.
Comment by miki123211 5 hours ago
Where the new generation of LLMs (Fable, Sol) shines is tasks that are much harder than typical soft eng, yet that still have a verifiable answer, think mathematical proofs or exploits. I think there's still a good amount of low-hanging fruit in those (and similar) areas.
The next frontier after that is tasks that don't have automatically-verifiable answers, and may not even have correct and incorrect ones in the strictest sense of the word.
Reasonable lawyers might disagree on the question of "which trial strategy do I use given the following set of facts." There are answers that are clearly wrong, but being able to choose between many plausibly-correct ones requires many years of lawyering and seeing many trials play out. I do suspect that most lawyers are far below the ceiling that a hypothetical immortal lawyer that has practiced for an infinite amount of time would have achieved.
Comment by sscaryterry 6 hours ago
Comment by rllearneratwork 7 hours ago
Comment by Zambyte 7 hours ago
Comment by bitexploder 7 hours ago
One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.
Comment by monster_truck 6 hours ago
Setting the memory to "fast timings" is good for 8-12% more tokens/second if you haven't tried yet. I miss the slightly older days of AMD when powerplay tables were unlocked and we could configure the timings and voltages manually, there's another 30% being left on the table ez
Comment by MrDrMcCoy 1 hour ago
Comment by snapplebobapple 7 hours ago
Comment by icedrift 8 hours ago
Comment by jimbo808 7 hours ago
Comment by dw_arthur 3 hours ago
Comment by d2p 9 hours ago
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
Comment by h14h 9 hours ago
https://artificialanalysis.ai/methodology/intelligence-bench...
Edit to provide AA's article explaining it:
https://artificialanalysis.ai/articles/artificial-analysis-i...
Comment by kmeh 6 hours ago
Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.
Comment by nolok 6 hours ago
Comment by ahartmetz 9 hours ago
Comment by splatzone 7 hours ago
Comment by johnnyApplePRNG 8 hours ago
Comment by torginus 7 hours ago
Comment by gpt5 8 hours ago
Comment by Gcam 5 hours ago
The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.
Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...
Comment by saretup 4 hours ago
Comment by TacticalCoder 4 hours ago
Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?
That's indeed a bit fishy.
Comment by personjerry 8 hours ago
Comment by apitman 8 hours ago
Comment by WD-42 9 hours ago
Comment by eli 10 hours ago
I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
Comment by thefourthchime 9 hours ago
Comment by ghosty141 6 hours ago
Comment by sscaryterry 6 hours ago
Comment by comboy 9 hours ago
Comment by eli 9 hours ago
OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.
Comment by trey-jones 9 hours ago
Comment by Gooblebrai 9 hours ago
Comment by lkt 7 hours ago
Comment by Gooblebrai 6 hours ago
Comment by iAMkenough 9 hours ago
Comment by MrDrMcCoy 1 hour ago
Comment by bitexploder 7 hours ago
Comment by g58892881 9 hours ago
Comment by onomojo 9 hours ago
Comment by cromka 9 hours ago
Comment by cromka 7 hours ago
Especially the second one seems exactly like my experience.
Comment by hungryhobbit 8 hours ago
It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?
Comment by kloop 3 hours ago
Comment by moffkalast 8 hours ago
With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or maybe just thinks for 10 minutes instead, then fixes one thing and breaks four additional ones.
Comment by msp26 8 hours ago
But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.
Comment by cromka 7 hours ago
Comment by msp26 7 hours ago
"give me this again without jargon invented this session at high density
and with a couple (maybe more or less) simple useful ascii diagrams underneath each design"
The context is that I was discussing an experimental new idea for my video game review analysis product.
Designs 1,2, and 3 were horrible: the model even suggested a rejection after the word soup so it would have been pointless to waste my fleeting time on Earth reading it.
Otherwise, I generally really enjoyed using fable for bouncing ideas. It was an absolute joy to have this thing provide useful criticism, analyse sample data, and create prototypes so that I could elevate my understanding of the problem without stepping down from a pure intuition/design headspace.
But I don't consider the purely model written code usable for a feature this important. I'll probably scrap it entirely and start from scratch with newfound understanding.
Comment by sscaryterry 6 hours ago
Comment by copperx 9 hours ago
Comment by garciasn 9 hours ago
Comment by aenis 9 hours ago
I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.
Comment by petesergeant 9 hours ago
Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.
Comment by hbn 8 hours ago
I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.
I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.
Comment by moffkalast 8 hours ago
Comment by sscaryterry 6 hours ago
Comment by usef- 9 hours ago
Comment by efficax 7 hours ago
Comment by PacificSpecific 3 hours ago
Comment by cromka 8 hours ago
Comment by yeeeloit 3 hours ago
Comment by nimonian 8 hours ago
Comment by TacticalCoder 7 hours ago
To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.
Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".
As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.
Comment by visarga 9 hours ago
Comment by capnjazz 9 hours ago
Comment by FridgeSeal 7 hours ago
“One thing worth your attention, if you were to detonate a pipe bomb in your house, it would have a negative effect on your living room”.
Comment by ethin 5 hours ago
Comment by greenchair 8 hours ago
Comment by vunderba 6 hours ago
It will do everything it can to defer or push it off, to the point where I’ve had to add multiple imperative directives to the AGENTS file telling it, in no uncertain terms, not to defer tasks under any circumstances.
Comment by cyanydeez 5 hours ago
Comment by vunderba 5 hours ago
• Qwen3-VL picks up new images in a NAS, auto captions and adds the text descriptions as a hidden EXIF layer into the image, which is used for fast search and organization in conjunction with a Qdrant vector database.
• Gemma3:27b is used for personal translation work (mostly English and Chinese).
• Some small 8b models (like llama3.1) for sentiment analysis on text.
But haven't really tried using local LLMs in conjunction with agentic harnesses yet.
Comment by cyanydeez 5 hours ago
My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context. opencode's dynamic context pruning plugin can get you pretty far into the stratosphere.
Comment by smartbit 2 hours ago
See https://news.ycombinator.com/item?id=48883538 25 days ago
> The Sleev (the project has been renamed to make a startup) creator was shilling their project in the OpenCode Discord. That person is very convinced they have something that no one has ever built before. They focused on token reduction without any real evals for capability impacts.
I'm generally against this context pruning without prompting or details. Sleev is very opaque about how it works and definitely will bust your cache.
Comment by vunderba 5 hours ago
Thanks for the tip - I like this a lot. I remember having to do a lot of tweaking to curtail Qwen QwQ-32b when it would go down an endless psychotic recursive reasoning loops as part of its "chain of reasoning."
Comment by conception 1 hour ago
I don’t see anyone talking about how you have to completely change your prompting strategies with Op. 5 versus 4.8 to get the most success.
Comment by Fordec 8 hours ago
Comment by thomasfromcdnjs 7 hours ago
I could not get Opus 5 to do anything without losing a few years of my life from stress.
Fable has been okay but I am doing ML work and not allowed to use it which feels insane.
Comment by CuriouslyC 8 hours ago
Comment by combyn8tor 7 hours ago
Comment by enraged_camel 9 hours ago
After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.
My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.
Comment by cesarvarela 8 hours ago
Comment by fellowniusmonk 8 hours ago
Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.
Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.
Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.
Comment by nomel 9 hours ago
Comment by dgellow 7 hours ago
Comment by ofjcihen 6 hours ago
Comment by drschwabe 8 hours ago
Comment by kachnuv_ocasek 8 hours ago
Comment by petesergeant 9 hours ago
Comment by logicchains 9 hours ago
Comment by pornel 9 hours ago
Comment by paradox460 2 hours ago
Comment by dr_dshiv 8 hours ago
Comment by vardalab 7 hours ago
Comment by sunaookami 8 hours ago
Comment by bontaq 9 hours ago
Comment by seizethecheese 9 hours ago
Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol
Source: http://pellmell.ai/leaderboard.
This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
Comment by d4rkp4ttern 8 hours ago
Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in plain terms.
For a fairly gnarly task, after fighting with with Claude-Code + Opus-5, I ported my session to Codex + GPT-5.6-sol, and it was like a breath of fresh air.
Arguably a key aspect of intelligence is concise, clear communication, and current benchmarks miss that, at least as far as I'm aware. I would think some arena-type benchmarks where humans rate responses would measure this, though I'm not sure which those are.
Comment by chpatrick 8 hours ago
Comment by msp26 7 hours ago
Try asking them to make useful diagrams for some stuff in a codebase, out of the box without excessive hand holding they don't make good choices about what's worth communicating and how to do it.
You see this in their pointless frontend copy all the time too.
Comment by embedding-shape 6 hours ago
Same concept of "over-sharing" seems to prevalent in a bunch of domains when it comes to LLMs, sometimes more visible, sometimes less.
Comment by cindyllm 6 hours ago
Comment by gpt5 8 hours ago
Comment by chpatrick 7 hours ago
Comment by cloverich 7 hours ago
Comment by riknos314 7 hours ago
If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.
Comment by IanCal 8 hours ago
Comment by computably 8 hours ago
Comment by micw 8 hours ago
Comment by HappyPanacea 8 hours ago
Comment by a2ff6eeb0 8 hours ago
Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.
Comment by FridgeSeal 7 hours ago
Comment by veber-alex 7 hours ago
Comment by veber-alex 7 hours ago
This community is pure trash.
Comment by HDBaseT 5 hours ago
"You can adjust the output style in your '.claude/settings.local.json' file".
OR
"You can decrease verbosity by doing x, y and z."
Comment by jiggawatts 8 hours ago
I wonder if this is a side effect of MoE models — they can write excellent prose, but not simultaneously with writing code.
Comment by notfromhere 5 hours ago
And writing doesn’t have validators like code so you can’t really scale it in the same way
Comment by octoberfranklin 6 hours ago
Yeah if you ask it a medical question, it answers in that impenetrable jargony style that clinical journals use... full of unnecessarily custom adjectives ("orthopedic" instead of "of the bone") and discipline-specific terms (anterior, distal) even when the user didn't display mastry of this terminology (hint to frontier labs: add training cases for this; it will improve your model's usability).
My theory is that LLMs perceive the writing styles of various fields as being like different (but related) languages, and they're inclined to answer a question in the language of its source material unless specifically asked otherwise. If you add "ELI5" the model treats it as a question plus a translation task.
I think this is why programming questions are answered with an exaggerated cringey form of HN-speak ("load bearing", "gate" as a verb, "dissolves") by some models.
Comment by moffkalast 8 hours ago
Comment by fearmerchant 8 hours ago
Comment by pixelready 7 hours ago
Comment by satvikpendem 7 hours ago
Comment by hungryhobbit 7 hours ago
I mean, I get it: how fast a model responds is relevant. But a test that changes second by second is far less relevant than a test that tells you how smart the model is, and accounts for latency in some way that isn't constantly changing the result.
Comment by seizethecheese 7 hours ago
Latency and throughput matter a ton as a user, so I think it's actually totally defensible for a leaderboard to bounce around a lot as these numbers change. The best model to use changes a lot based on these!
Comment by embedding-shape 10 hours ago
Comment by artemisart 9 hours ago
Comment by moritzwarhier 9 hours ago
But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.
A lingering sense of quality degradation when going deep remains.
But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well.
And on-prem capability is simply irreplaceable.
Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.
Comment by amelius 9 hours ago
Comment by user43928 9 hours ago
$0.36 per task, Intelligence Index score 56 -> Grok 4.5 high
$1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max
$0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh
$1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh
Comment by scrlk 10 hours ago
> Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)
> Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested for the Coding Agent Index.
Comment by apitman 9 hours ago
Comment by Bootvis 9 hours ago
https://artificialanalysis.ai/models/qwen3-8-max
Doesn't have the claim either. Clickbait?
Comment by petu 9 hours ago
Comment by Bootvis 9 hours ago
Even then, this seems a much more marginal win than the headline suggested to me.
Comment by theropost 9 hours ago
Comment by tarnith 8 hours ago
I've had to use it a bit for work, and it's been remarkable watching the degradation in performance with the default suggested current models (Opus 5 as a prime example) vs the models that got them huge attention a year ago (Opus 4.6)
If you give 4.6 a spec, or existing code to implement a feature in, it will ask some pointed questions if there's something unclear in the spec, and then produce a plan and move to implement it.
5 will freak out at even a basic task, ask itself if it's own assumptions or your instructions are correct, proceed to re-assess it's own plan, and it's instructions 3-4 times, and then maybe produce code after burning several hundred thousand tokens (and quite a bit of time) analyzing existing code and thoroughly sweeping it for irrelevant problems both to the task it was given and the spec it came up with.
It's quite bizarre to me how well advertised the benchmarks and anecdotes from people one shotting MVP browser games are, compared to the experience of everyone I know that's had to actually use it to accomplish even a relatively basic task.
Comment by aenis 9 hours ago
Comment by gnull 8 hours ago
That vibe coding they brag about as if it was a good thing, it shows.
Take their notation for describing permissions. The docs are not comprehensive, and in practice it doesn't quite work how they describe it.
Or their management of sub-agents. I once lost a sub-agent, it finished and disappeared from UI. Apparently, you can't bring it back yourself: you have to ask the parent agent to do it for you. But the parent was Fable, and I ran out of credits, so I was locked out of using my opus sub-agent because of it.
Or an even more grotesque example: when you paste your claude API token to authorize, it covers characters with *. But it seems like an LLM has hallucinated a limit of API key length and the tail of your key stays visible.
Comment by hungryhobbit 8 hours ago
I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment.
IF YOU ARE GOING TO HAVE A SHITTY VIBE CODED PRODUCT, AT LEAST USE YOUR SHITTY AI TO FIX THE SHITTY PROBLEMS!
Comment by cindyllm 7 hours ago
Comment by thejosh 8 hours ago
I can't believe how many critical bugs fall through.
My favourite one is the bug where Plan mode can execute destructive commands inadvertently.
Then all these get closed with `Closing for now — inactive for too long. Please open a new issue if this is still relevant.`. Awesome.
Comment by ethin 5 hours ago
Comment by formerly_proven 8 hours ago
Almost like CC is 100% vibe coded.
Comment by tempest_ 9 hours ago
Opus 5 is just a token burner.
I use fable plan and spawn opus 4.8 workflows which seems to work alright.
Comment by ethin 5 hours ago
Of course, the hilarious thing to me is that Anthropic likes to claim that the usage limits are because of resource allocation problems or something like that. Obviously no such issue exists, otherwise they wouldn't allow you to bypass it by just paying a bit more and it would be a hard limit. So usage credits are entirely their way of just screwing you out of more money.
Comment by aenis 8 hours ago
For me it is not great for design work - Fable is way better, and 4.8 was conservative and thus better (Opus 5 seems to jump to conclusions far more eagerly). But for overnight builds, where I give it 8hrs worth of work on LLDs created by Fable - its great. Where Opus 4.8 would often lose the plot and stop for questions clearly answered in the LLD - Opus 5 does manage to complete. Since it launched, I don't remember it ever disappointing me with builds. But designs? Boy, is this thing explosively stupid sometimes.
Comment by robbru 8 hours ago
Comment by mikae1 9 hours ago
Comment by arrowleaf 8 hours ago
Comment by bakugo 8 hours ago
Comment by enedil 5 hours ago
Comment by arikrahman 8 hours ago
Comment by tyre 8 hours ago
Comment by gamblor956 8 hours ago
Comment by swalsh 8 hours ago
Comment by CuriouslyC 8 hours ago
Comment by cortesoft 9 hours ago
With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.
Comment by notatoad 7 hours ago
Comment by AlexandrB 8 hours ago
Comment by riknos314 7 hours ago
Comment by ericd 8 hours ago
Comment by dexwiz 8 hours ago
Comment by ericd 6 hours ago
I don't think you're right about that last prediction, at all. And new use cases are opening up as these get smarter. I think things are going to get pretty weird.
But the point was that it really doesn't look like they're losing money on users, on average.
Comment by swalsh 8 hours ago
Comment by hahahaa 8 hours ago
There is some element of responsibility on the user to guide and monitor the model/harness and not let it rip to burn tokens.
Comment by criddell 8 hours ago
Maybe they consider that hiring a person to do it would have cost at least as much and taken much more time, so paying them is a bargain.
Comment by echelon 8 hours ago
Plus we get to own, keep, run, do whatever with the model. We don't feel trapped. Moreover, it's something we can truly build on top of and own our own destiny.
Anthropic and OpenAI are the new Oracle (Oracle pre-AI; Oracle is even worse now). Expensive, feels like dealing with a lawyer, and not at all open. They just became infinitely less cool than they were a month ago.
The whole of our industry is going to migrate to open weights. We're smart enough to know this is the better deal and technical enough to be able to pull it off.
The only thing that might save these OpenAI and Anthropic in the near-term is an abundance of enterprise contracts negotiated with non-tech companies. They'll soak consulting firms and F500 companies for "AI" integrations.
Comment by criddell 8 hours ago
I think that's exactly what they are going for - enterprise and government customers.
Comment by pvtmert 8 hours ago
Amazon's first principle is the Customer Obsession. Making customers happy.
Fun bit is that the human psychology rates personal looking fixes better than having no issues at all.
For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment".
Meanwhile, any other (small) cloud. Simple, no weird charges. Even _most_ of network egress is free. But, no reason to call support or feel "extraordinary". Comes out as "meh" against Amazon's "top tier" support model...
Comment by john01dav 8 hours ago
Comment by riknos314 6 hours ago
Anthropic trains models on AWS's (and GCPs, and Microslop's) infrastructure, then skims margin off of selling inference also on the infrastructure owned by the other companies.
These are extremely different businesses.
Comment by axpy906 7 hours ago
Comment by polishdude20 9 hours ago
Comment by petercooper 9 hours ago
Comment by quotemstr 9 hours ago
People produce such models by over-RL-ing smaller models on math and coding tasks. I've found the results capable of neither innovative work nor thinking outside the box. They're straight-A students raised by tiger moments who never let them play freely for hours in the dirt.
Perhaps you could say such models are skilled --- but intelligent? Not by my measure.
People and AIs alike need diversity of experience and a broad liberal arts education to see hidden connections between fields and make real advances.
Comment by DC-3 9 hours ago
Comment by petercooper 8 hours ago
Sticking to LLMs, they seemingly get their intelligence (whatever that really means) from building models rich with knowledge, so you could have a point. But Qwen models seem to be particularly good, even at small model sizes, at maintaining both their own knowledge while acquiescing to and integrating external information in the moment.
Comment by syntaxing 10 hours ago
Comment by colingauvin 9 hours ago
Comment by tarr11 9 hours ago
Comment by syntaxing 8 hours ago
Comment by LoganDark 9 hours ago
Comment by markasoftware 9 hours ago
Comment by dofm 8 hours ago
But if you are sort of pair-programming with the model, the speed obviously matters and I think then the 35B is acceptably smart, and when it's wrong it'll be wrong much more quickly. It seems very good on SQL and PHP, and I assume on typical JS and Python.
I would rather work that way, so I hope they do produce a small MoE model.
Comment by CamperBob2 9 hours ago
Comment by 13rac1 9 hours ago
Comment by syntaxing 9 hours ago
Comment by SwellJoe 10 hours ago
I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.
It is smart and reasonably quick but not reliable.
Comment by superfrank 9 hours ago
At their best, I think they're closing in on Opus and GPT, but they're incredibly inconsistent and the variance in output quality is much higher than the best from any of the Anthropic or OpenAI models from the last few generations. The only way I can describe it is that it feels like a lack of intuition with the models which means I find my self needing to write longer prompts or have more back and forth to get them to do what I want from them.
To give an example, I have a saved prompt that I use as a sanity check on some data I'm storing. It reads about 50 rows from a DB and matches them to the UI and makes sure the data is displaying correctly. I've been using this with GPT 5.5 and now 5.6 for a few months and running it a few times a week with no issue. Sometimes I'll run it multiple times in a single chat if I notice bad data (run it, fix thing, run again, fix another thing).
I recently tried to switch to using Deepseek v4 (first flash and then pro) and while both did the task just fine, both would do things like change the response format from one message to another in the same chat or randomly decide to omit things it didn't think were relevant. At one point I ran the prompt, fixed some bad data, and then said "Okay, I fixed row 7, run {prompt} again" and so it decided to leave row 7 out of the response. A few times the first message would contain a table and then the next run in the same chat would contain the data in a bulleted list.
None of those are major issues and all could be solved with a bit more rigor in my prompting, but for me it makes them harder to work with. Those examples are a bit trivial, I think they're the easiest way for me to illustrate the gaps I see with them.
Comment by dyauspitr 9 hours ago
Comment by gerdesj 4 hours ago
Qwen3.6-27B-FP8 currently espouses (see below), which is not too bad. The directions are a bit mad but the mileage is about right and there is a helicopter manufacturer here and a RNAS (navy not airforce) museum nearby at Yeovilton. Cosford is in Shropshire which is not a million miles away.
I'm not sure what 盆地的 means but the river Yeo is correct ... OK ... "basin like" - again not bad, even if Chinese is not the first language here. The model understands that Yeovil is named after (or vice versa or at least is associated with) a river
Yeovil is the current form of Gifle (Saxon) which I thought meant "bend in a river" but WP is currently saying "fork in a river". My source is a local museum. There is a fork but was it there 2000 odd years ago? My hydrology skills say ... possibly
---------------------------------------------------- Q: where is yeovil:
Yeovil is a town in Somerset, in the South West of England.
It is located roughly:
25 miles (40 km) south-west of Exeter
60 miles (100 km) west of Bristol
140 miles (225 km) west-south-west of London
Yeovil is known for its historic market town center, RAF Museum Cosford (nearby), and as a significant industrial town, particularly during World War II for aircraft manufacturing (including the Wellington bomber). It sits in the盆地的 valley of the River Yeo.Comment by quirino 9 hours ago
I wasn't able to find an explanation from them. Anyone knows what happened?
Comment by Art9681 9 hours ago
Comment by ignoramous 9 hours ago
Comment by drnick1 10 hours ago
Comment by eli 10 hours ago
Many providers will host it and will compete on price. It also can't easily be taken away because one company (or one government) decides they don't want it around any more. People can fine-tune it for particular workloads.
Comment by Art9681 9 hours ago
Might as well use gpt-sol.
Comment by iAMkenough 8 hours ago
I stopped paying attention to self-published benchmarks when Apple started using those non-sensical performance graphs with "relative performance" as a vertical axis when announcing a new chip.
Comment by drnick1 9 hours ago
It's barely better, and barely cheaper, not really enough to challenge the status quo IMO. Half the price for basically the same performance would be a much stronger value proposition.
Comment by ux266478 9 hours ago
Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.
Comment by SwellJoe 8 hours ago
Also, OpenRouter misses most of the usage of the US models, as most people are getting those from the vendor directly via subscriptions.
Comment by eli 8 hours ago
Comment by ux266478 8 hours ago
You can argue there's a selection bias that openrouter users are less likely to display model loyalty, but it would still be a visible confounding factor if it was a statistically significant behavior. And it's not. Nor is there a visibly meaningful indication that people don't sidegrade between models. With every single data set, you're going to see that. You're also going to see it reflected in discourse, as I mentioned. Fact of the matter is there isn't a status quo in AI any more than there's a status quo in cars.
Comment by apitman 10 hours ago
Comment by frereubu 9 hours ago
Comment by apitman 9 hours ago
That said, it's a fair point. For me, it boils down to things covered here: https://earendil.com/posts/session-portability/
Things like obscured reasoning traces.
Comment by copperx 9 hours ago
Comment by jjice 9 hours ago
Comment by benjiro29 8 hours ago
What cost the most in API. Input, Cached Input, or Output. There you have your answer.
Unfortunately, we have moved so much of the actual intelligence of models towards reasoning, what results in some models getting good scores, but this is because they are dumping a insane amount of reasoning tokens at the problem.
So a mid priced model, with heavy reasoning output, cost the same as a expensive model, with medium reasoning output.
Before the GPT Luna price drop of 80%, you actually had the same price if you used Luna High and Sol Low. With the difference that Sol Low was insane fast, and often way better code.
Do not look at the top score but more what is on the horizontal axis as you go down. Sol Medium is frankly, was the best performance for dollar, until that Luna price drop. I will even argue that despite the higher price, Sol Medium is still way better despite Luna Max being cheaper. Or Opus Low, one of the better values also.
What do you notice? Is that those models all have a high intelligence start point for their low setting. So that means they do not rely as much on output tokens aka thinking.
Comment by criley2 9 hours ago
There are a couple of frontiers (ok bad word, maybe categories) in open weight models.
These Qwen 3.8 and Kimi K3 style models aren't trying to win on price, they're trying to compete on intelligence and capability.
Models like Deepseek V4 Flash (updated this week) are $0.03 a task, or 50X cheaper than Qwen3.8/Kimi K3, and 100X cheaper than Fable, while offering stunning intelligence. That's a different frontier for competition, and perhaps one more interesting for someone who wants to see them compete on cost.
Comment by ecocentrik 9 hours ago
Comment by Alpha3031 9 hours ago
Comment by jazzyjackson 9 hours ago
Comment by TheCycoONE 9 hours ago
Comment by efficax 9 hours ago
Comment by ngl999 1 hour ago
Comment by MrDrMcCoy 1 hour ago
Comment by aliljet 10 hours ago
Comment by Alpha3031 9 hours ago
Comment by teravor 9 hours ago
Comment by zmmmmm 7 hours ago
Comment by mindwok 6 hours ago
Comment by h14h 9 hours ago
Comment by bonoboTP 8 hours ago
Comment by ben8bit 9 hours ago
Comment by tomComb 9 hours ago
I was with you until there. Qwen and the OpenAI models are great, aggressive agents, but they’re not as good as the anthropic models for human interaction. They just don’t have the subtlety, understanding, or attention to detail.
Comment by colingauvin 7 hours ago
Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy. I'm constantly fighting with it to stop fighting me and accept that I occasionally know better. It's really frustrating to spend so many tokens of such an expensive model arguing with it.
Comment by ben8bit 9 hours ago
Comment by londons_explore 8 hours ago
Clearly the weighting of those things depends on the usecase
Comment by camnora 9 hours ago
Comment by brettgo1 8 hours ago
Comment by daemonologist 8 hours ago
Comment by arjie 8 hours ago
$25k - DSv4 Flash
$4k - Qwen 3.6 35A3B Q5
$1k - Qwen 3.6 27B Q4
Some people prefer the sense over the MoE YMMV.
Comment by apitman 6 hours ago
Comment by Fordec 8 hours ago
Comment by dangoodmanUT 7 hours ago
Comment by steve-atx-7600 9 hours ago
Comment by looksjjhg 9 hours ago
Comment by Footprint0521 5 hours ago
Comment by brcmthrowaway 9 hours ago
Comment by colingauvin 9 hours ago
[0]https://humanparadox.org/local-vs-frontier-benchmarks-for-my... - note here I tested Q8 but have found no difference at lower quant.
Comment by dofm 7 hours ago
Comment by LPisGood 9 hours ago
Comment by kyxsc 8 hours ago
Comment by notatoad 6 hours ago
Comment by delduca 10 hours ago
Comment by sirbor 9 hours ago
Comment by esafak 8 hours ago
Comment by atemerev 9 hours ago
Comment by ramon156 8 hours ago
Comment by dyauspitr 9 hours ago
Comment by proxyscore 8 hours ago
Comment by OsamaJaber 8 hours ago
Comment by indiantrains 9 hours ago