56 points piotrgrabowski 1 hour ago 22 comments
snarfy 47 minutes ago | parent
simonw 47 minutes ago | parent
Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
M3 Ultra 31 23.5 20 15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...peri-cl 14 minutes ago | parent
Also: ~30 token/s on GLM 5.3-flash, locally.
/meta Here's a CSS filter that stops those nuisance chart animations,
macstories.net##*:style(animation: none !important; transition: none !important)redox99 13 minutes ago | parent
peri-cl 6 minutes ago | parent
gpugreg 12 minutes ago | parent
RationPhantoms 12 minutes ago | parent
Maybe Apple is an acquisition away from changing that balance.
sajithdilshan 43 minutes ago | parent
That’s like 12 years worth of OpenAI Pro subscriptions
geodel 34 minutes ago | parent
Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.
vardump 32 minutes ago | parent
Kurtz79 6 minutes ago | parent
A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.
simonw 31 minutes ago | parent
Plenty of other reasons to get excited about it local AI, but I don't think cost is one of them.
criddell 6 minutes ago | parent
And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.
112233 24 minutes ago | parent
ApolloFortyNine 37 minutes ago | parent
I didn't expect this to make the 5090 to look like a good deal.
kokonokko1337 36 minutes ago | parent
Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.
srcreigh 32 minutes ago | parent
I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.
It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.
The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.
An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
WarmWash 31 minutes ago | parent
Ehh, the actual elephant in the room is:
"why bother with local AI at all when you can lease a GPU for $5/hr?"
To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.
tempoponet 27 minutes ago | parent
This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.