The Rack Is Not Inevitable
A large language model lives in a rack — rows of accelerators, vast pools of high-bandwidth memory, another data center and another power contract with every increase in size. That sequence has become so familiar it sounds like physics. But it is not physics. It is an architecture.
What if a machine Apple shipped in 2019 could run a 744-billion-parameter model — not by squeezing it all into memory, but by learning to move the right pieces at the right time? This essay follows that experiment and what it says about where intelligence is allowed to live.

