M5 Ultra Mac Studio Review
Posted by piotrgrabowski 1 day ago
Comments
Comment by simonw 1 day ago
Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
M3 Ultra 31 23.5 20 15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...Comment by gpugreg 1 day ago
Comment by beastman82 1 day ago
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
Comment by _hugerobots_ 1 day ago
Comment by louthy 1 day ago
Perhaps consider some non-offensive language for your comparison?
Comment by pseudosaid 19 hours ago
Comment by louthy 15 hours ago
Ok, as somebody with ADHD I find it offensive because I don't need constant supervision, implying people with ADHD need constant supervision is belittling and just plain wrong. So, I will call out an offensive trope if I see it.
> If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith
No, because if there was some truth to it, I wouldn't be offended. Perhaps stop with the amateur psychology? You're not very good at it.
Comment by nacs 1 day ago
Comment by ProllyInfamous 1 day ago
My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."
This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
Comment by selectodude 1 day ago
Comment by tomega2134 1 day ago
Comment by throwaway219450 1 day ago
32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.
Comment by searealist 22 hours ago
Comment by throwaway219450 19 hours ago
Comment by searealist 17 hours ago
Comment by throwaway219450 2 hours ago
Comment by fhub 1 day ago
Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.
Comment by throwaway27448 1 day ago
Comment by bigyabai 1 day ago
No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.
Comment by throwaway27448 1 day ago
Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.
Comment by bigyabai 1 day ago
But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.
Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.
Comment by girvo 1 day ago
Comment by Eisenstein 1 day ago
Comment by girvo 1 day ago
Comment by medvezhenok 1 day ago
Comment by spider-mario 9 hours ago
Comment by beastman82 1 day ago
Comment by cyanydeez 1 day ago
Comment by mathisfun123 1 day ago
Comment by throwaway27448 1 day ago
Comment by mathisfun123 1 day ago
Comment by throwaway27448 1 day ago
I don't get these weird parasocial emotional attachments/beefs people have with brands. Talk to a therapist.
Comment by mathisfun123 1 day ago
Comment by selectodude 1 day ago
Comment by tom_ 1 day ago
Comment by _hugerobots_ 1 day ago
Comment by bigyabai 1 day ago
Comment by _hugerobots_ 1 day ago
Comment by bigyabai 1 day ago
Comment by liuliu 1 day ago
Comment by GeekyBear 1 day ago
Comment by searealist 1 day ago
Comment by ActorNightly 1 day ago
Comment by bee_rider 1 day ago
Comment by abletonlive 1 day ago
Comment by ActorNightly 1 day ago
Comment by abletonlive 1 day ago
Well, if those are the only two options you can come up with it's pretty clear that this isn't about me or what I am, you have a false model of reality.
> Running very large models on Mac is unusable at 10 tok/sec.
There are plenty of examples of models running at well over 10 tok/sec that aren't viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?
I think you're projecting pretty hard with the two options you've listed. Go touch some grass, you seem overly frustrated that reality doesn't meet your expectations.
Comment by ActorNightly 23 hours ago
Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.
Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper
The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.
So I ask you again, which one are you?
Comment by abletonlive 20 hours ago
No thanks, you're not in a position to do that clearly.
> Since you clearly don't use local llms
I do, probably a lot longer than you have actually.
> anything under 100 tok/sec is USELESS
Objectively wrong. You sound like you're really behind and you're so myopic that you think coding is the only use case for local LLMs. I'm a professional software dev and that's the least interesting use case of local LLMs.
> Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec.
You clearly didn't read the article or have reading comprehension issues. The model is Qwen3.8-Flash-Next 4 and 5-bit quant, neither of which "conveniently fits on one GPU". Sorry that your hardware doesn't live up to your own delusions and can't even run Qwen3.8-Flash-Next at 4/5 bit quant. You are taking the Quen3.8-27B numbers, something that the article isn't really that concerned with, and trying to make it fit into your narrative.
> So I ask you again, which one are you?
Well I'm someone that suggests that you should touch some grass and reevaluate your personal issues. You seem angry. Perhaps it's best to figure your own issues before trying to figure out why people are excited about Apple hardware for local llms. I am sure the people that need to interact with you in society would be very grateful if you took the time to do this.
Comment by ActorNightly 17 hours ago
A) He literally says "I tested a different Qwen model for the comparisons between Mac and PC." The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this https://ollama.com/library/qwen3.8:27b. And the speed of that model on a 5090 in terms of tok/sec is not 60 lol.
B) If M5 ultra runs 40 tok/sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: https://ollama.com/library/qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok/sec lol.
As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out https://huggingface.co/Qwen/Qwen3.8-27B versus https://qwen.ai/blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models running dogshit slow at peak ~20 tok/sec?
It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.
Considering that Apple probably doesn't want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.
Comment by peri-cl 1 day ago
Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think).
/meta Here's a CSS filter that stops those nuisance chart animations,
macstories.net##*:style(animation: none !important; transition: none !important)Comment by redox99 1 day ago
Comment by tcdent 1 day ago
Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.
Comment by nojs 1 day ago
Inference time is going to be dominated by the low memory bandwidth on these Macs, so a dense model will suffer most. It’s more of an opportunity for large MoE models with a low number of active experts since you can keep all experts in VRAM but not pay the bandwidth cost until they are used.
> you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers
This is an interesting direction that I expect to see more of. But for most models currently you need basically all experts loaded since they are chosen per token.
Apple seems to be researching longer horizon expert caching, where they keep experts swapped in for longer runs of tokens [1]. Other labs are offloading ngram caches but not sure if they’re pursuing anything like this?
1. https://machinelearning.apple.com/research/introducing-third...
Comment by thejazzman 20 hours ago
Comment by peri-cl 1 day ago
Comment by skohan 1 day ago
Comment by nacs 1 day ago
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
Comment by peri-cl 1 day ago
https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...
(Note it's a sparse MoE with only 6B active).
Comment by bitexploder 1 day ago
I paid $500 for the RAM in Nov 2023 :)
Comment by peri-cl 1 day ago
No wonder Warren Buffet gave up and resigned.
Comment by bitexploder 1 day ago
Comment by nacs 1 day ago
That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
Comment by well_ackshually 1 day ago
Comment by nacs 1 day ago
If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).
Comment by cma 22 hours ago
Comment by karmakaze 1 day ago
These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.
Comment by smcleod 12 hours ago
Comment by lhl 1 day ago
This is with llama.cpp. You can of course use vLLM/SGLang well on these cards and they're even faster. On vLLM w/ NVIDIA/Qwen3.8-27B-NVFP4 baseline has a prefill of about 13,000 tok/s. The baseline tok/s is 72 tok/s, but at mtp7, it's 157 tok/s, and w/ dflash7 that goes up to 215 tok/s. On mtp-bench, DFlash2 gets a hair under 300 tok/s w/ the code_python prompt.
Comment by alex7o 1 day ago
Comment by RationPhantoms 1 day ago
Maybe Apple is an acquisition away from changing that balance.
Comment by wlesieutre 1 day ago
Comment by kridsdale1 1 day ago
Comment by dagmx 1 day ago
Apple just shifted to N2. They’re not going to be doing another major shift right away.
And TSMCs own roadmap would put your hallucination years away at best for a a product that follows a roughly annual cadence https://www.tomshardware.com/tech-industry/semiconductors/ts...
Comment by smith7018 1 day ago
[1] https://wccftech.com/apple-to-move-to-1-4nm-process-soon-to-...
Comment by GeekyBear 1 day ago
> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators
https://www.tomshardware.com/tech-industry/semiconductors/ap...
Comment by aurareturn 1 day ago
Comment by api 1 day ago
Comment by jmyeet 1 day ago
This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.
I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.
Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.
The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.
Comment by pama 1 day ago
Comment by glitchc 1 day ago
Comment by throw0101c 1 day ago
Why Infiniband ("IB")? If it's for RDMA, that is possible with certain Ethernet cards/chipsets as well. Certainly Mellanox, but Broadcom:
* https://techdocs.broadcom.com/us/en/storage-and-ethernet-con...
and Intel as well:
* https://www.intel.com/content/www/us/en/support/articles/000...
Link level flow control or priority flow control needs to be supported on the switch ports as well.
Comment by wmf 1 day ago
Comment by fragmede 1 day ago
Comment by happyopossum 1 day ago
Where are you buying 8 5090s for under $10k? With CPU, RAM, and (checks comment) infiniband hardware???
You're probably looking at a lot closer to $60k when all is said and done, and that's before you hire an electrician to run a sub panel for your homelab...
Comment by kridsdale1 1 day ago
Comment by prmoustache 1 day ago
Comment by jmyeet 1 day ago
Each PC is probably going to cost ~$6k and you're talking about 8000W of electricity draw. That's going to consume multiple 20A circuits even at 240V. And the electricity ain't free either. A Mac Studio seems to draw ~500W max.
Oh and the Mac Studio has an upgrade route to run 1T+ models too by chaining them together with TB5 chaining. OSX supports RDMA this way. That's comparable bandwidth to the 100Gbps Infiniband option.
So you're talking about $50-60k of hardware and more power draw and more heat for something that will I'm sure beat the MS M5U option but at huge cost. Also, at that kind of price point, I'm likely to get a workstation PC and put 2 (or possibly 3) 6000 Pros in it.
Comment by weee322 1 day ago
every company make his own npu (without xai)
probaby in 2028 we will have more concurent firm on market place
Comment by poisonxa 7 hours ago
Comment by traceroute66 1 day ago
Comment by washadjeffmad 1 day ago
nvidia-smi -pl 450 for like a 4% reduction in throughput. I tend to set it around 350W because it's a comfortable temperature blowing on my legs under the desk without warming my office in the summer.
I put together this system two years ago, so it's a little out of date, but it only cost $3000 for the same performance and capability as an Ultra. I don't think I would spend $7000 to save 100W, though.
Comment by TacticalCoder 1 day ago
Yeah people don't pay enough attention to those settings IMO. The first thing I do when I set up a new machine (or upgrade my OS) is to restore all my powersaving configs.
For example I've got all but one of my virtual desktops that put the CPU in powersave mode: I don't need max Ghz when browsing the Web, not even on demand. But when I switch to the virtual desktop where my development environment is, then I want power on demand.
Now I don't do it to save the planet: I do it because I love a quieter computing experience (coupled with Be Quiet! PSU and Noctua fans, this makes for a very quiet computer). That it consumes less electricity is a nice side-benefit.
Comment by beastman82 1 day ago
Comment by ActorNightly 1 day ago
Comment by GeekyBear 1 day ago
> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman...
Apple plans to release a base M6 chip this fall for entry-level Macs... a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028.
https://www.tomshardware.com/tech-industry/semiconductors/ap...
Comment by srcreigh 1 day ago
I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.
It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.
The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.
An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
Comment by zozbot234 1 day ago
Comment by srcreigh 1 day ago
isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.
> but this would decrease single-session performance even further
Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]
It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.
And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see
If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.
If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.
[1]: https://omlx.ai/benchmarks/performance/69kzkrv8?utm_source=c...
Comment by zozbot234 1 day ago
Comment by slowin 1 day ago
Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.
That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).
Comment by nowittyusername 1 day ago
Comment by Octoth0rpe 1 day ago
I think this is true, but also misses that a lot of us are just doing basic flask apps with a react front end. We don't need astra; Something sonnet 4.6 level locally is perfectly sufficient 95% of the time, and maybe 99% of the time.
Comment by brandon272 1 day ago
It's like watching a discussion about cars available to take on a 100km road trip. A new car gets released that is on par with a Toyota Corolla but it is dismissed as completely useless for a 100km trip because it doesn't have the seat massagers and air ride suspension that the new Escalades have.
The reality is that something like Sonnet 4.6 is still amazingly capable for so many programming tasks, especially if you already have some reasonable level of experience to steer it in the right direction.
And if you think Sonnet 4.6 is still worthwhile, then it seems undeniable that something like Qwen 3.8-27B is also worthwhile.
Comment by dash2 1 day ago
Comment by _hugerobots_ 1 day ago
Comment by slowin 1 day ago
Comment by _hugerobots_ 1 day ago
Comment by slowin 1 day ago
I'm also a huge fan of local models and think it's absolutely imperative that they continue to advance so we can move off of the Anthropic/OpenAI hosted models. It's important to accurately asses where we are in that journey though.
Comment by srcreigh 1 day ago
Comment by slowin 1 day ago
Comment by _hugerobots_ 1 day ago
Comment by fhub 1 day ago
Correctness matters much more than speed to me, but if I can get both, that’s obviously very interesting.
Comment by brandon272 1 day ago
Comment by slowin 1 day ago
Comment by brandon272 1 day ago
Comment by slowin 1 day ago
Comment by brandon272 1 day ago
Comment by slowin 23 hours ago
Comment by poincareball 8 hours ago
Comment by sajithdilshan 1 day ago
That’s like 12 years worth of OpenAI Pro subscriptions
Comment by 112233 1 day ago
Comment by woah 1 day ago
Comment by techmunky 1 day ago
Comment by glitchc 1 day ago
Do they include footguns from pointer bugs?
Comment by Razengan 1 day ago
Comment by 112233 18 hours ago
https://youtu.be/AQf84KubJjE https://youtu.be/RAUsCwu8ekQ https://youtu.be/NtsJ5m6C7dU https://youtu.be/h8xO4PiJ--w
Comment by Razengan 15 hours ago
The internet is still alive
Comment by kridsdale1 1 day ago
I appreciate the reference to RUSH: Red Barchetta in the final line.
Comment by nowittyusername 1 day ago
Comment by Octoth0rpe 1 day ago
Comment by zamadatix 1 day ago
Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.
Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".
Comment by cma 22 hours ago
Not necessarily for MoE
Comment by throw0101c 1 day ago
I think most people are getting 512 for running Chrome with a bunch of tabs open. /s
Comment by geodel 1 day ago
Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.
Comment by jeffybefffy519 1 hour ago
Comment by Kurtz79 1 day ago
A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.
Comment by qwytw 1 day ago
Is there evidence that's true though? I mean gross margins on subscriptions being negative since the API is seemingly very profitable (if the price is compared with the cost of serving very large open models).
As long as there is pressure from other providers serving cheaper models that are somewhat competitive without having to incur any of the R&D costs raising prices will be tricky.
Comment by BatFastard 1 day ago
Don't you mean an Apple to NVidea comparison?
Comment by vardump 1 day ago
Comment by cyclopeanutopia 1 day ago
Comment by prmoustache 1 day ago
Comment by patrickmcnamara 1 day ago
Comment by geodel 1 day ago
I think it goes without saying. And it is eminently evident over last couple of decades that from compute to storage to meals 3rd part providers have saved billions upon billions of dollars to enterprises and individuals alike by providing these essential services.
Comment by kridsdale1 1 day ago
Comment by geodel 1 day ago
Comment by ericmay 1 day ago
[1] https://www.macworld.com/article/3238319/mac-studio-m5-max-r...
Comment by simonw 1 day ago
Plenty of other reasons to get excited about local AI, but I don't think cost is one of them.
Comment by criddell 1 day ago
And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.
Comment by simonw 1 day ago
Comment by bel8 1 day ago
It's a bold strategy cotton, lets see if it pays off for em.
Comment by auntienomen 18 hours ago
Comment by Danox 1 day ago
Comment by hgoel 1 day ago
Comment by Danox 1 day ago
Comment by ionwake 1 day ago
Comment by SXX 1 day ago
Might be if RAM prices get much more reasonable its gonna be 1/3 of the price, but it's very much possible its gonna be half or more.
And if you're buing Mac Studio and not some AI-only locked down board it's possible to reuse it for other purposes.
Comment by matt-p 1 day ago
Comment by Danox 1 day ago
Comment by tempoponet 1 day ago
This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.
Comment by _hugerobots_ 1 day ago
Comment by ApolloFortyNine 1 day ago
I didn't expect this to make the 5090 to look like a good deal.
Comment by nacs 1 day ago
It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.
Comment by orsorna 1 day ago
Comment by peri-cl 1 day ago
(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).
Comment by orsorna 1 day ago
Comment by asimovDev 1 day ago
Comment by Eisenstein 1 day ago
Comment by asimovDev 12 hours ago
Comment by liuliu 1 day ago
Comment by kokonokko1337 1 day ago
Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.
Comment by Danox 23 hours ago
Why waste time spending years reverse engineering. Wouldn’t it have been easier to raise money for such an endeavor? I hope someone will come along in the future in kindergarten, junior high or high school who doesn’t know any better will try something like this. I’m too old. for such a journey.
Comment by steve1977 1 day ago
I get it on Windows systems, at least when someone wants to use Linux-type tooling. But macOS already supports pretty much all of that natively?
Comment by dylan604 1 day ago
Comment by RunSet 1 day ago
For starters, the source code.
Comment by steve1977 1 day ago
Apart from that, for the UNIX part, the source is available for quite a few components:
https://github.com/apple-oss-distributions
notably also the kernel
Comment by throw0101c 1 day ago
Strictly speaking, Apple can claim to ship a UNIX® operating system:
Comment by Gracana 1 day ago
kokonokko1337 already said it was good enough to run LLMs, presumably RunSet isn't saying the source code is needed to run an inference server.
Comment by odkdkekfkwjf 1 day ago
Comment by hamiltont 1 day ago
I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
Comment by peri-cl 1 day ago
For generation speed in isolation, yes.
Comment by GeekyBear 1 day ago
Comment by akozak 1 day ago
Comment by saagarjha 1 day ago
Comment by happyopossum 1 day ago
Comment by Danox 22 hours ago
The Mac Pro Tower plus the Cinema Display monitor at the time was even more, similar to the Studio M5 ULtra today than the iMac which I believe was an upper middle computer?
$5,700 in 2011 is worth $8,488.46 today
$7,700 in 2011 is worth $11,466.87 today
The cost of Mac Studio M5 Ultra today $9,466.87-$11,466.87
Note: If memory cost was the same as it was two years ago subtract about $2000-$3000 dollars.
Today’s prices are not that far off from top end computers.
Comment by sethd 1 day ago
I ordered the same one for work so I could run more local agents at once (many iOS simulators and Xcode build processes).
Comment by mtsolitary 1 day ago
Comment by SamuelAdams 1 day ago
Comment by jjtheblunt 1 day ago
Comment by Gracana 1 day ago
If I switch to Mac OS, I have to sort out a package manager and install all the stuff that's missing, and when it comes to containers... they're just linux VMs. I'd happily cut out the weird proprietary middleman if I could.
Comment by Danox 1 day ago
Comment by jjtheblunt 4 hours ago
Comment by Gracana 1 hour ago
Comment by crossroadsguy 1 day ago
Comment by mjlee 1 day ago
I'd be quite surprised if Mac OS alone needs more than 8GB, given that they sell the Neo with 8GB of RAM today.
Comment by odkdkekfkwjf 1 day ago
Comment by theplumber 1 day ago
Comment by lowbloodsugar 1 day ago
Comment by mstaoru 1 day ago
Comment by manyatoms 1 day ago
You run uncensored local models where you can ask questions that would get denied by public providers, or questions that you prefer them not to know the intricate details (like your financial planning)
Comment by mstaoru 16 hours ago
Comment by apexalpha 15 hours ago
There's value in knowing what capabilities local models will have on regular hardware in the near future.
Comment by EricE 23 hours ago
Comment by addaon 1 day ago
Comment by snarfy 1 day ago
Comment by andrekandre 1 day ago
but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)
Comment by drdaeman 1 day ago
Comment by bel8 1 day ago
A $10/mo subscription to OpenCode Go would have done the job for you.
They have models like Kimi K3, Grok 4.6 , GLM-5.3, Mimo 2.6 Pro (launched today, already available) which are happy to follow your orders without accusing you of being a terrorist.
Comment by chasd00 1 day ago
Comment by drdaeman 1 day ago
I thought this only applies to LLM inference providers, but not raw GPU rentals.
Comment by manyatoms 5 hours ago
Comment by Eisenstein 1 day ago
Comment by BatchJob 1 day ago
Comment by crorella 1 day ago
Comment by devy 1 day ago
Comment by aenis 1 day ago
Entry level serious hardware starts at 100k, and a bit better but still almost-useful grade is 200k (8x rtx pro, plus a nice epyc pairing). Thats the sort of thing a salaried expert lets their employer buy them for sort of serious work.
Anything really serious is well north of 1M - not including the housing and commercial grade mains connection. And at best that buys fast Kimi K3 or GLM.
Comment by apexalpha 15 hours ago
These models are not 'barely capable', they're contemporary near frontier.
Comment by aenis 12 hours ago
Comment by apexalpha 11 hours ago
Comment by 12kaj2 1 day ago
Comment by prmoustache 1 day ago
Comment by villgax 1 day ago
Comment by saejox 1 day ago
Comment by vivzkestrel 20 hours ago
- https://sunkcost.ai/s/mac-studio-m5-max-128/qwen3.8-27b-q4/?...
Comment by slashtom 1 day ago
Comment by sghiassy 1 day ago
Comment by whalesalad 1 day ago
Comment by jmull 1 day ago
99% of people will use whatever AI is free. The sophisticated, heavy users that are willing and able to pay a lot of money the ones that will be interested in controlling their inference bills.
Today, the sweet spot where an M5 Ultra makes sense is tiny. But we might expect that to grow a lot.
Comment by BatFastard 1 day ago
Even if you could get a frontier model, you would not be able to run it on any Mac. So speculating on what M7 or M9 will achieve in 5 years (if we even still exist) seems pointless.
Comment by sghiassy 1 day ago
I don’t think Apple is going to lie down and cede AI to the cloud.
Comment by geodel 1 day ago
How about writing mail to President and senators on AI doomsday scenario if frontier labs do not pace themselves?
That mini model on mac mini would scared to hell to do such thing. It need that rugged frontier model to speak truth to power.
Comment by BatFastard 1 day ago
I would love to find an excuse to buy a 10,000 dollar machine! But I cant find one yet. My current cloud bill is in excess of 400 USD per month. Just can't achieve frontier model capabilities locally.
Comment by Danox 23 hours ago
Comment by sghiassy 1 day ago
Comment by whalesalad 1 day ago
Comment by sghiassy 1 day ago
Comment by fragmede 1 day ago
Comment by beastman82 1 day ago
Comment by CamperBob2 1 day ago
Sam's address will probably be more riveting, imaginative, and terrifying than the last couple of Terminator screenplays. Legislators will lobby him to write the laws for them, and the ghost of Harlan Ellison will threaten to sue him.
Comment by ajross 1 day ago
I really don't see who buys this, except people who want the Studio for some other reason. But nothing in the story says you want to fill racks with these instead of Blackwell or TPU parts; it's not even close.
Comment by sghiassy 1 day ago
Think of a company like Apple moving onto your turf. They’re not going to cede AI to the cloud. They want their part of the pie.
So in 7 years, how much AI will be handled locally on your iPhone. And will you have repaid all the debt on your balance sheet before Apple eats your lunch
Comment by SXX 1 day ago
It's way too easy for 1T+ frontier labs to ditch Nvidia. So Nvidia will also put effort to make sure there are open weights models and local hardware available.
And Apple will benefit from this too.
Comment by ajross 1 day ago
Comment by cptskippy 1 day ago
However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there's a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.
Comment by Danox 23 hours ago
Those companies building those big data centers are going to find out that they overspent. Because we’re not going back to the mainframe era, no matter how much OpenAI, Anthropic, Microsoft, Meta or Google would like to.
Comment by lowbloodsugar 1 day ago
Comment by rwissinger 1 day ago
Comment by itsmeduncan 1 day ago
Comment by WarmWash 1 day ago
Ehh, the actual elephant in the room is:
"why bother with local AI at all when you can lease a GPU for $5/hr?"
To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.
Comment by Youden 1 day ago
Unless you only need the AI available some of the time, $5/hr is pretty expensive. That's an RTX Pro twice a year.
If you're using it for discrete sessions of coding or something, that might make sense for you but if you're using it for an always-on assistant, that pricing kinda sucks.
Comment by WarmWash 1 day ago
5090's are like $0.20/hr
Comment by brandon272 20 hours ago
Comment by chasd00 1 day ago
there are lots of people with very expensive hobbies, see sailboat racing for example.