Apple is positioning the new Mac Studio as a “local AI desktop.” This third instalment looks at what its memory options mean in practice and whether owners of an M3 Ultra should upgrade.
Why memory matters: A local language model has to fit in memory while it runs. With the commonly used 4-bit quantisation for open-weight models, one billion parameters take roughly half a gigabyte. On that rough basis, a 96 GB entry configuration can comfortably run models around 70 billion parameters and may be tight at 120 billion; 256 GB reaches roughly 400 billion, and 512 GB around 600 billion by our calculation. A secondary report cited by Memeburn says Apple points to models around 400 billion parameters for the 512 GB configuration. If a model does not fit, it may not load or may become impractically slow when parts are read from storage.
The second factor is memory bandwidth. As a model generates each word, it repeatedly reads weights from memory, so output speed roughly depends on bandwidth relative to model size. The M5 Ultra’s 1.2 TB/s is 46% higher than the M3 Ultra’s 819 GB/s; the M5 Max’s 614 GB/s is about half the Ultra’s. Apple’s “up to four times faster prompt processing” claim in LM Studio refers to processing a long input prompt, not generating output, and is attributed to new Neural Accelerators in the GPUs.
Clustering: New Mac Studios can connect over Thunderbolt 5 using RDMA and share a model. Apple says four machines provide inference up to three times faster than one. This may be an alternative for people who cannot wait for or afford 512 GB, but no independent measurement has been published yet.
Should M3 Ultra owners upgrade? For video editing and general work, probably not: the reported 30% CPU and 60% GPU gains do not clearly justify a new machine costing around 370,000 TL, and the current model remains powerful. For local language models, it depends: if a model cannot fit on a 96 GB M3 Ultra, more memory may be the only route; Apple removed the M3 Ultra’s 512 GB option this year amid memory constraints. For AI training and image generation, the upgrade could make sense if Apple’s claimed 3.3× to 4.3× gains hold, but waiting for independent confirmation is prudent. The 512 GB models are not expected before late October.
The series’ conclusion: the M5 Ultra story is about capacity and AI more than raw speed. It is a solid generation-on-generation update for conventional work; for local models, it could be a major shift if Apple’s claims prove accurate. We will update this series when independent AI benchmarks are published.
Verified: Bandwidth figures, clustering and the 3× claim, 512 GB timing, and reduction of the M3 Ultra configuration: Macworld and 9to5Mac. The 4× LM Studio figure is an Apple claim reported by 9to5Mac. Uncertainty: Memory-per-parameter is a rule of thumb and varies by model and quantisation. The 400-billion figure comes from a secondary source (Memeburn). The upgrade recommendation is our analysis. Not available: Independent tokens-per-second measurements for M5 Ultra language models have not been published.

Leave a comment