This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.
I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.
It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.
The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.
An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
show comments
sajithdilshan
On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.
That’s like 12 years worth of OpenAI Pro subscriptions
show comments
sethd
I find it funny that the thing always mentioned with this machine is local AI. If you're a local model enthusiast, then maybe that makes sense, but I just don't see the economics working there.
I ordered the same one for work so I could run more local agents at once (many iOS simulators and Xcode build processes).
ApolloFortyNine
The model being tested is 18k as configured.
I didn't expect this to make the 5090 to look like a good deal.
show comments
tempoponet
While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.
This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.
show comments
hamiltont
Once you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB/s. IMO best ROI for most people is 'cheapest used Ultra with enough RAM'
I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
show comments
mstaoru
What do people realistically do with these? It's too slow and du... not SOTA-level for coding. It's way too slow for video. I tried simulating an "Fable herding Qwen subagents" and it takes much longer and delivers a much worse result than Fable/Astra alone.
show comments
akozak
"a total cost of $0" Uhh ... how much is that hardware?
show comments
liuliu
When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
mtsolitary
Waiting for my 64GB M5 Pro Mini, hoping it will also be fun to tinker with for local AI
SamuelAdams
I think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.
show comments
crossroadsguy
My mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That's a shame. If only I could install an alternative OS that uses very little amount of RAM :-)
show comments
kokonokko1337
> "It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"
Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.
show comments
theplumber
At this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing
show comments
addaon
Ordered one for OpenFOAM. Excited for it. Will be nice to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.
snarfy
$12,299
show comments
crorella
What are good options to run local models nowadays? Something good for coding and personal assistant kind of things
BatchJob
If you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.
devy
This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!
show comments
12kaj2
The Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.
show comments
saejox
i can buy a house with that amount of money. it used to be car money.
villgax
Lol, try generation of images & videos on these, they ought to improve perf on Deep learning not just llms
slashtom
Fantastic review, this is how it should be done with local AI.
sghiassy
Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI
show comments
cptskippy
I think we'll eventually get to the point where folks will have a local AI agent but I think people need to temper their expectations to a degree. You aren't going to have data center level tok/s from a box sitting under your desk and you don't need instantaneous responses for many workloads. Having a local agent that can execute tasks over a couple days with your supervision that might otherwise take you weeks is perfectly acceptable.
However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there's a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.
lowbloodsugar
Everyone looking at the Qwen3 27B model and the 5090. It’s like saying a Porsche is better than a $5m Komatsu earth mover at moving a 20lb carry-on suitcase. Yes. Yes it is. Why do people spend $5m on a komatsu then, when this one metric shows the Porsche is better? Huh. Show us the “Moving 300 metric tonnes in one load” metric. How’s the Porsche now? Oh, the Porsche is in the Komatsu? Ok I’m getting a bit carried away with that analogy.
WarmWash
>Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?
Ehh, the actual elephant in the room is:
"why bother with local AI at all when you can lease a GPU for $5/hr?"
To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.
I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.
It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.
The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.
An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.
That’s like 12 years worth of OpenAI Pro subscriptions
I find it funny that the thing always mentioned with this machine is local AI. If you're a local model enthusiast, then maybe that makes sense, but I just don't see the economics working there.
I ordered the same one for work so I could run more local agents at once (many iOS simulators and Xcode build processes).
The model being tested is 18k as configured.
I didn't expect this to make the 5090 to look like a good deal.
While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.
This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.
Once you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB/s. IMO best ROI for most people is 'cheapest used Ultra with enough RAM'
I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
What do people realistically do with these? It's too slow and du... not SOTA-level for coding. It's way too slow for video. I tried simulating an "Fable herding Qwen subagents" and it takes much longer and delivers a much worse result than Fable/Astra alone.
"a total cost of $0" Uhh ... how much is that hardware?
When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
Waiting for my 64GB M5 Pro Mini, hoping it will also be fun to tinker with for local AI
I think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.
My mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That's a shame. If only I could install an alternative OS that uses very little amount of RAM :-)
> "It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"
Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.
At this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing
Ordered one for OpenFOAM. Excited for it. Will be nice to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.
$12,299
What are good options to run local models nowadays? Something good for coding and personal assistant kind of things
If you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.
This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!
The Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.
i can buy a house with that amount of money. it used to be car money.
Lol, try generation of images & videos on these, they ought to improve perf on Deep learning not just llms
Fantastic review, this is how it should be done with local AI.
Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI
I think we'll eventually get to the point where folks will have a local AI agent but I think people need to temper their expectations to a degree. You aren't going to have data center level tok/s from a box sitting under your desk and you don't need instantaneous responses for many workloads. Having a local agent that can execute tasks over a couple days with your supervision that might otherwise take you weeks is perfectly acceptable.
However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there's a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.
Everyone looking at the Qwen3 27B model and the 5090. It’s like saying a Porsche is better than a $5m Komatsu earth mover at moving a 20lb carry-on suitcase. Yes. Yes it is. Why do people spend $5m on a komatsu then, when this one metric shows the Porsche is better? Huh. Show us the “Moving 300 metric tonnes in one load” metric. How’s the Porsche now? Oh, the Porsche is in the Komatsu? Ok I’m getting a bit carried away with that analogy.
>Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?
Ehh, the actual elephant in the room is:
"why bother with local AI at all when you can lease a GPU for $5/hr?"
To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.