Salesforce Global Outage

Posted by mabil 3 hours ago

Counter149Comment78OpenOriginal

Comments

Comment by stmw 41 minutes ago

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

Comment by shimman 20 minutes ago

This is the case for every single B2B saas product. This is like the "bar is rolling on the floor" level of competence required. Please have higher standards for paid products.

Comment by mulmen 24 minutes ago

I honestly don’t get the snark. The status page has:

Seemingly meaningful IDs

Search

Region filter

Email update signup

Predictable URLs for instance status so they can be deep linked in runbooks

What appears to be the actual live instance status.

What appears to be the actual live service status in each instance.

An update log with frequent detailed updates.

Comment by dominotw 3 minutes ago

> I honestly don’t get the snark.

I get it. Fat fuck didnt show an ounce of grace to anyone his whole life so why should we return the favor. Has to be one the most outwardly vile human being after larry ellison

Comment by raffraffraff 2 hours ago

Have you tried turning it off and then on again?

> We're no longer pursuing restarts as a path to remediation.

Oh you have

Comment by cube00 1 hour ago

Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it.

> We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.

At least it didn't fix the problem so they can actually start finding the real cause.

> We're no longer pursuing restarts as a path to remediation.

Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?

Comment by swatcoder 11 minutes ago

I don't know, that reads exactly like an AI troubleshooter working through a plan without the implicit contextual understanding an experienced human might bring to either the actions or the communications.

"Oops, we forgot to tell it that this is the hyperscaled Salesforce production environment and that its choices need to project competence and consider brand embarrassment. WILLFIX"

Comment by jakevoytko 16 minutes ago

In my experience it’s a safe way to do something useful while everyone is getting their bearings. It immediately partitions the situation space between being persisted or systemic vs local or caused by long-running processes. Plus everyone’s going to ask if you’ve tried that already, so you might as well get it out of the way if it makes any amount of sense

Comment by andrewinardeer 1 hour ago

"Yeah, I'm with Rob. Just let's reboot and see what happens"

Comment by chihuahua 9 minutes ago

"If that doesn't work, clear the cache and reboot again."

Comment by 1 hour ago

Comment by mergy 1 hour ago

Unplanned outage timing is never good but this is really not good.

https://www.salesforce.com/dreamforce/

Sept 15-17

Comment by ramesh31 36 minutes ago

Probably not a coincidence

Comment by minimeow 33 minutes ago

This is what happens when more than half the company is away attending the Salesforce cult-indoctrination stuff while spending all their bandwidth making customers/partners feel good.... The stuff that matters to keep the lights on gets overlooked.

Comment by gigatree 31 minutes ago

Arguably, making customers and partners feel good is the more important part of the business

Comment by walt_grata 20 minutes ago

They wont feel good if the product they pay for doesnt work

Comment by paimapi 18 minutes ago

arguably, this is what sales cares about and the half-measures taken to tackle what must be the Mt. Everest of tech debt at Salesforce is what leads to large, systemically degraded customer trust in products that keep shipping bugs

Comment by orochimaaru 6 minutes ago

I don’t think engineering and SRE of the organizing company are ever invited to those events. They’re mainly for marketing and sales (which includes solution architects).

Comment by dominotw 3 minutes ago

this event is only for customers. its not a company event.

Comment by fhub 1 hour ago

Cause: Legacy Salesforce login service got into a resource-exhaustion cascade.

Fix: Rolling some unspecified fix they proved in testing out over the fleet seemingly very slowly (After their earlier attempts to roll something out faster failed).

Details at https://status.salesforce.com/incidents/20004433

Comment by dd8601fn 6 minutes ago

I wonder if “legacy login” is the shared login gateway.

It’s optional but everyone uses it. And it was flaky for an hour or so, like two months ago.

Comment by cyberpunk 3 hours ago

That status page is the most salesforce thing ever.

Scroll down. >_<

Comment by sunrunner 2 hours ago

Wow, it's almost as long as the Every UUID V4 or Every Floating Point Number pages.

Comment by dgellow 3 hours ago

I see tabs with a spinner loading infinitely, which is indeed very salesforce like

Comment by troupo 2 hours ago

It lists all instances, you can drill down into each one of them, see which services are affected, and each affected service pops up the incident timeline?

Isn't it actually amazing, and not "the most salesforce thing ever"?

Comment by chrisjj 2 hours ago

Yup. No mention of outage. Even drilling down gets nothing more than "Service Disruption".

Comment by santoshalper 3 hours ago

At least they're consistent about UX. My only complaint is that it needs more tabs.

Comment by chasd00 2 hours ago

You can click on any of the instances and then the service that is down to read the updates. It’s not 100% clear but some sort of issue with a “legacy login service”. The latest updates say a fix is rolling out.

Comment by electroweak 1 hour ago

At this point OpenAI really ought to let us know when they're testing again.

Comment by AmazingTurtle 2 hours ago

You better bet someone started their agents with a prompt "Make a salesforce clone but with 100% uptime"

Comment by Bluestein 2 hours ago

"Make a salesforce clone but with 100% uptime"

  ⎿  You've hit your session limit · resets 2:53am (48°52.6′S, 123°23.6′W Etc/GMT+8)
  /upgrade to increase your usage limit.

Comment by arionmiles 3 hours ago

This is certainly a unique status page.

Comment by danjc 2 hours ago

It's dns isn't it

Comment by dogas 10 minutes ago

I was scrolling to see this comment, haha

Comment by bearjaws 1 hour ago

Feel like it has to be for all of this to go down at the same time.

Comment by Bluestein 2 hours ago

It's always DNS.-

Comment by ajross 1 hour ago

More like 70% human-configured DNS, 25% human-configured routing configuration, 5% interesting software bug.

Comment by Bluestein 1 hour ago

Entirely correct.-

(Nowadays any of those need to fit in an "agent dropped all tables. Apologized" moment.-)

Comment by ghusto 2 hours ago

No, IPv6

Comment by afr0ck 2 hours ago

Classic

Comment by alienbaby 2 hours ago

Haven't they got some kind of new fancy ai interface they can use to fix it?

Comment by adamw2k 3 hours ago

Perfect timing with Dreamforce this week.

Comment by ownerr 2 hours ago

Here is link to incident details: https://status.salesforce.com/incidents/20004433

Comment by ssk42 2 hours ago

My gut instinct is that this is about when all of their on prem servers were EOL and their /public cloud solution was required. This must have had something to do with that

Comment by BoorishBears 2 hours ago

My gut instinct is that this is about Dreamforce with the rickshaws and whatnot

Comment by bgro 49 minutes ago

Leetcode developers win again

Comment by lrvick 1 hour ago

Seems it is back up now. Damn.

Comment by sidcool 40 minutes ago

ClaudeForce in action!

Comment by cmiles8 2 hours ago

Dreamforce, wake up. You’ve overslept!

Comment by dogas 9 minutes ago

Comment by mariopt 3 hours ago

Could it be people doing Claude/GPT automations and they just can't handle it?

Comment by mulmen 32 minutes ago

Could be. With a big conference going on and the updates mentioning resource exhaustion it could be a bunch of people doing demos, AI driven or not. Basically slashdotting themselves.

Comment by voidUpdate 2 hours ago

Wow, the intern must have tripped over a very big power cable this time

Comment by Havoc 2 hours ago

So I guess today everyone gets actual work done

Comment by monkeydust 2 hours ago

Sam speaks at Salesforce.

Salesforce goes down.

No causation here...move on.

Comment by hgoel 1 hour ago

Wasn't it Dario

Comment by jtrn 2 hours ago

And nothing of value was lost. God, I hate everything about Salesforce. Sometimes I have to integrate against their services, and it is always a pain, not to mention what the core project actually is: optimization of marketing and spam.

Comment by chasd00 33 minutes ago

not really defending salesforce but OAuth+REST is a pain? Pretty plain vanilla in terms of integration requirements.

Comment by jackdecker 2 hours ago

Ah yes. Exactly what a status page should look like: an endless list of random ID’s that don’t mean anything and no information whatsoever

At least salesforce is consistent with their design language

Comment by dd8601fn 13 minutes ago

Random Ids? If you mean the “USA324” ones, those are pods. If you’re a customer you know which one(s) you care about.

Comment by chasd00 2 hours ago

> random ID’s that don’t mean anything and no information whatsoever

If you use salesforce you know what all of that stuff means. Just click on one, it’s not rocket surgery.

Comment by jackdecker 2 hours ago

Was really just poking fun at them - AWS’s status page isn’t much better

Comment by 2 hours ago

Comment by BoorishBears 2 hours ago

Looks like a region list to me, maybe just with a lot of regions

Comment by 2 hours ago

Comment by 2 hours ago

Comment by aeneas_ory 2 hours ago

It’s akways DNS or login

Comment by theshrike79 3 hours ago

Was it DNS? Any guesses? :)

Comment by dude250711 1 hour ago

Vibe-slop, push to Prod?

Comment by geerlingguy 1 hour ago

They let their agentic AI take over system maintenance at Dreamforce yesterday like they were pushing in the talks

/s (partly)

Comment by sparkling 3 hours ago

Remind me please, what are folks currently paying per seat for this glorified CRUD app?

Comment by warmedcookie 2 hours ago

It's this thing called golf course driven development

Comment by electroweak 44 minutes ago

Permanent cache for that one.

Comment by raverbashing 2 hours ago

If I had been just out of Uni I would think this is edgy

Now I'm just glad I'm not responsible for this fire

Comment by themgt 2 hours ago

Never before in the history of global compute outages was so little lost by so many down servers, whose purpose was known to so few.

Comment by elzbardico 15 minutes ago

Kind of ironic. Salesforce is basically one of the major spiritual grandfathers of Slop. It is not uncommon in production systems to find that objects like Contact and Account have hundreds of custom fields. Sometimes, you find out that several of them have the same meaning and semantics, but were used at different times. Digging out you discover that some Marketing guy that used to work at the company did some task in a certain way that was lost when he was gone, and then a few months later his substitute had the same need and went ahead and created the same field with a slightly different name.

Doing data engineering work with Salesforce data is an exercise on archeology, psychology and organizational politics.

Slop is basically the ontological and teleological philosophy behind Salesforce very existence. Despite the official discourse that the "No Software" meant no infrastructure, no toil with updates and configuration, the subtext as intended for executives was very clear: "No need for you to be blocked by those pricks from engineering and their stupid, bureaucratic and gatekeeping rules".

"No software" was a call-to-arms to a certain subset of managers that were radicalized by Nicholas Carr's 2023 HBR article "IT Doesn't matter". It doesn't matter that Carr was a journalist and a writer with a masters in English that has never ever run even a small bodega, or has never managed an IT department. Anti-intellectualism and the abundance of capital brought in by the petrodollar that allowed the US government to run deficits year by year while exporting the ensuing inflationary effects to rest of world, would ensure that this message would ressonate and then even be amplified during the years of ZIRP and the Baillouts. Play fast and loose, first come, first served, a rising tide rises all boats and all that jazz. Wall Street favors bold, and the heck with the long term! This quarter will only live once!

Frankly, this is just poetic justice: Kill by slop, be killed by slop.

Comment by schnevets 2 minutes ago

Something tells me Troy the Salesforce Admin/BD Analyst did not cause the SAAS infrastructure to go down.

And I think you're confusing crud with slop.

Comment by danvesma 2 hours ago

"Trust just got personal!"

Comment by dboreham 3 hours ago

At least now we can figure out what Salesforce does.

Comment by sunrunner 2 hours ago

"And in tonight's news, the worldwide CRM solution Salesforce had a global outage affecting one hundred percent of its customer base. We interviewed users of the service to find out the scope of the impact. Everyone agreed that they were impacted, but strangely, nobody could describe _in what way_ they were affected."

Comment by justplay 20 minutes ago

[dead]

Comment by WhenItRuns 1 hour ago

[flagged]

Comment by irregularbowels 1 hour ago

[dead]

Comment by sick_of_slop 2 hours ago

[dead]

Comment by fr2029 3 hours ago

[dead]

Comment by gjvc 3 hours ago

[flagged]

Comment by mannycalavera42 2 hours ago

[flagged]

Comment by alienbaby 2 hours ago

Status page is still showing all red tho