Elevated errors for multiple models – Resolved
Posted by corvad 1 day ago
Comments
Comment by mariocesar 1 day ago
Comment by jerpint 1 day ago
The entire platform is skill driven, and based on the premise that state is your local file system. That makes switching harnesses so easy
It’s all open source and has plenty of other features, including inter agent communication, telegram client and much more in the pipeline
Comment by calgoo 21 hours ago
Comment by ludwik 17 hours ago
Comment by _flux 14 hours ago
Comment by etoxin 22 hours ago
Comment by vasco 20 hours ago
Comment by ammario 22 hours ago
Comment by adinb 18 hours ago
Comment by RGS1811 1 day ago
Comment by oulu2006 1 day ago
DS 4.1 flash is my main powerhouse and Opus/Astra my auditors (when they're not out of tokens) otherwise K3 or DS4 pro
Comment by SOLAR_FIELDS 1 day ago
Coming from mostly using Claude models, the terse factual statements coming from Astra via the Pi harness are a breath of fresh air over having to wade through the flowery verbose nonsense that Claude constantly outputs
Comment by pdntspa 21 hours ago
Astra found a number of flaws that would have come up during implementation and we worked through them. But then I had Fable 5.1 review that document and it found a number of issues with Astra's changes, the least of which had was that Astra duplicated a lot of technical notions that it added rather than using references to an authoritative section. It also flagged some of Astra's designs as technically impossible, pointing out why and I'm actually in the process of digesting its feedback and updating the design spec. (I hand-review each point and we work through a solution together -- I don't trust either model to come up with something that follows my vision on their own)
I'm not promoting one or the other, I just found it interesting how this sort of adversarial review found pretty significant flaws in the other model's work. I am curious as to whether this process will eventually converge on a document that both agree on or if the models are going to perpetually nitpick each other.
I haven't actually started implementation yet, so maybe one or the other is full of shit. Just trying to come up with an architecturally sound design for something I want to write, when I lack the DSP knowledge to be able to write it myself. But the intent is to pass an agent the design doc and list of milestones and let it handle implementation.
Comment by cgio 17 hours ago
Comment by pdntspa 10 hours ago
Comment by consumer451 1 day ago
1. who hosts the inference
2. which harness are you using with it, still CC?
Comment by oulu2006 1 day ago
1. I go direct to source, i.e. DS platform, I find it cheaper than paying the openrouter tax -- I also switch it up a bit
2. I built a local LLM router, that I update with new profiles that have my preferred provider of the week (lowest token costs/speed) with fallbacks, like mimo --> DS4 etc.. if there is overloading,
3. I use 3 diff harnesses, CC/Codex + Opencode -- they all talk to each other through a custom rig system that routes messages between llms using a Rust backed structured JSON system
Not saying this is the best, it's just what I like and works for me^.
I can flow quite naturally between Opus/Astra/K3/GLM/MiMo/DS/etc.. this way and often do...more so these days with subs no longer great as they used to be.
Comment by r_lee 15 hours ago
Comment by pavo-etc 1 day ago
Comment by RGS1811 12 hours ago
Comment by logicchains 20 hours ago
Comment by contentkraft 16 hours ago
Comment by someguyiguess 15 hours ago
Comment by mariocesar 15 hours ago
Comment by ebbi 1 day ago
Comment by drewnick 1 day ago
Comment by consumer451 22 hours ago
Comment by mariocesar 1 day ago
I also have the habit of naming my sessions with `/rename`
Comment by simlevesque 1 day ago
Comment by prodigycorp 1 day ago
gpt-6-sol and Aeon (personal agent) on Thursday. Already preceded by a huge week with step, mimo, grok, and jev releases.
Relentless cycle.
Comment by average_r_user 18 hours ago
Meanwhile, Meta's MUSE seems to be gaining traction in the US, while those of us in the EU are once again left watching from the sidelines.
Comment by bdcravens 13 hours ago
Comment by Razengan 1 day ago
Comment by prodigycorp 1 day ago
Comment by system2 1 day ago
Comment by AnotherGoodName 1 day ago
Comment by XenophileJKO 17 hours ago
Comment by Razengan 13 hours ago
Comment by shepherdjerred 1 day ago
Comment by Baeocystin 20 hours ago
Comment by Razengan 20 hours ago
In 1999 they even made a famous documentary about people in trench coats fighting AI
Comment by tombert 1 day ago
I am quite confident that what I'm doing is well within the law, and I'm not even doing any kind of pen-testing stuff, just some basic reverse engineering, but I can't even use Fable anymore because every time I enable it, it works for about twenty seconds and makes me drop down to Opus 4.8, and often even down to Sonnet.
If anyone here works at Anthropic, did you make the safeguards super sensitive recently?
Comment by theophilus76 1 day ago
Comment by chrisdbanks 1 day ago
Comment by blitzar 22 hours ago
Comment by Hamuko 22 hours ago
Comment by Schlagbohrer 16 hours ago
*if you live in a real country, which gives all workers unlimited paid sick leave, I am defining "max possible" here as "as many as you can reasonably take without your boss doing something about it"
Comment by Hamuko 12 hours ago
Comment by nikcub 1 day ago
Comment by Schiendelman 1 day ago
Comment by prologic 1 day ago
Comment by ehnto 1 day ago
HN is already a waterfall of AI meta conversations and bike-shedding, now we have to discuss service outages about the AI too?
Can we talk about stuff people are building again, with or without AI, and stop gasping at every minute detail of LLM service providers.
Comment by bdcravens 13 hours ago
Why are Apple product announcements upvoted? We've known about those changes for months, and they release on a predictable cadence.
Why did we ever care about San Francisco news? Most HNers aren't in SF, California, and many aren't even in the US.
etc
Comment by nullc 23 hours ago
Comment by system2 1 day ago
Comment by joegibbs 1 day ago
Comment by makeavish 1 day ago
Comment by Wowfunhappy 1 day ago
Comment by simlevesque 1 day ago
Comment by makeavish 10 hours ago
Comment by theGeatZhopa 19 hours ago
Comment by BoxOfRain 18 hours ago
Comment by camkego 1 day ago
Comment by homo__sapiens 1 day ago
Comment by bombcar 1 day ago
and keep workin'
Comment by TYPE_FASTER 1 day ago
Comment by bdangubic 1 day ago
Comment by butlike 10 hours ago
Comment by bonsai_spool 1 day ago
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
Comment by solenoid0937 16 hours ago
Comment by bbeonx 1 day ago
Comment by AnotherGoodName 1 day ago
Comment by a012 1 day ago
Comment by mattdeboard 1 day ago
Comment by porlex 1 day ago
Comment by itssohailkhan 15 hours ago
Comment by jcims 1 day ago
Comment by thenipper 1 day ago
Comment by 3ln00b 15 hours ago
Comment by yrcyrc 1 day ago
Comment by throw567643u8 1 day ago
Comment by windexh8er 1 day ago
Comment by Schiendelman 1 day ago
Comment by ArcHound 17 hours ago
We should hold corporations valued in billions to a higher standard than a single person trying to care for their family.
Comment by windexh8er 16 hours ago
I've recently started reading "End Times Fascism" [0] and it truly encompasses where we are at, how we got here and the bleak future we may be in for if we don't get in front of this now. But this book definitely highlights a lot of what some of the tech community has been seeing and hand waving about for years. The worn out trope that we should hold corporations accountable doesn't work. We should hold the people at the top accountable like other nations do because if they want the limelight then they should bear the responsibility that comes with abuses. The US is becoming a caste system by design. National surveillance, erosion of access to resources for personal compute, bought and paid for politicians and laws that don't apply to the wealthy.
[0] https://us.macmillan.com/books/9780374621384/endtimesfascism...
Comment by windexh8er 6 hours ago
Comment by ryanschaefer 15 hours ago
Comment by guybedo 1 day ago
Comment by addag 14 hours ago
Comment by mococa 13 hours ago
Comment by sharts 5 hours ago
Comment by NoPicklez 21 hours ago
Comment by joshtronic 1 day ago
Comment by calvinmorrison 1 day ago
Proof that vibe coded or not, people pay someone else mostly for liability.
Comment by corvad 1 day ago
Comment by Aeolun 1 day ago
Comment by bugfix 1 day ago
Comment by 15155 1 day ago
Comment by krembo 22 hours ago
Comment by rvz 1 day ago
If outages like this happened on every major deploy at any other company; this would be viewed as unacceptable, especially if it was something like Google Search going down on every update.
Comment by bofia 23 hours ago
Comment by acedTrex 1 day ago
Comment by dougame 22 hours ago
Comment by dragonlin 1 day ago
Comment by loverofspades 23 hours ago
Comment by iammrpayments 1 day ago
Not sure how can claude even compete at this point unless openAI seriously downgrade the models to save money.
Comment by hombre_fatal 17 hours ago
Comment by moecables 1 day ago
Comment by jakozaur 17 hours ago
We now have alien intelligence that is very different from human intelligence.
Comment by sick_of_slop 13 hours ago