- Sun 05 July 2026
- technology
- Gaige B. Paulsen
- #apple, #technology
How a 1984 Cray X-MP/48 actually stacks up against a decade of iPhone silicon — and why the honest numbers are smaller than the marketing math
There's a familiar genre of technology writing: take some storied old machine, hold it up against a modern phone, and marvel that the phone is thousands of times faster. The Cray X-MP/48 — the machine that rendered early Pixar work and stood in for the computers of Jurassic Park — is a favorite target. And the thousand-fold claims aren't wrong, exactly. They're just measuring the wrong thing.
If you compare the machines carefully — matching precision to precision, and looking hard at memory bandwidth, the axis where the Cray was genuinely world-class — a more interesting picture emerges. The supercomputer holds up far better than the headline numbers suggest. On the things it was actually built to do, the gap between 1984 and 2025 is closer to two orders of magnitude than four.
I have a personal reason for caring about this particular machine. In January 1986, I was working as an operator on NCSA's X-MP/48 at the University of Illinois, and I drew what remains one of the oddest assignments of my career: setting up the modem connection that let Arthur C. Clarke "christen" the machine from Sri Lanka. The christening was a typed message — Clarke, with his characteristic sense of occasion, evoking the birth of HAL 9000, who in 2001 becomes operational at a plant in Urbana, Illinois, a few minutes' walk from where the Cray actually sat. We unveiled it at an event at the Krannert Center. Standing there watching Clarke's words arrive down a phone line in the town HAL was born, I don't think any of us were doing arithmetic about FLOPS. But I've thought about that machine's numbers many times since — and they're not what most people assume.
Disclaimer: I want to note up front that I'd been contemplating writing this article for probably over a decade at this point. However, it wasn't until I could "collaborate" with modern-day AI that I had the time to get the research put together
The machine that was king
The Cray X-MP/48, in its canonical form, ran four processors at a 105 MHz clock (a 9.5-nanosecond cycle), reached a theoretical peak of 800 MFLOPS, and carried 64 MB of fast static main memory. It was the fastest computer in the world from 1983 to 1985.
Two things made it special, and neither was raw clock speed. First, it was a native 64-bit double-precision machine — every one of those 800 million floating-point operations per second was a full FP64 operation, the currency of serious scientific computing. Second, and more importantly, it had extraordinary memory bandwidth for its era. Each processor had four parallel paths to memory, and the memory could stream operands at clock speed. That's why real applications hit an unusually high fraction of peak: on the National Center for Atmospheric Research's production workload, the X-MP sustained about 225 MFLOPS — roughly 25% of theoretical peak, a remarkable figure when many machines of the day struggled to reach 10%.
A note on "peak." Every number in this article pairs a theoretical peak with a realistic sustained fraction. Peak FLOPS is what the transistors can do if perfectly fed; sustained is what real code actually gets. The gap between them is where all the interesting engineering lives — and, as we'll see, both the Cray and the iPhone hit surprisingly similar fractions of their respective peaks.
Four eras of iPhone
Rather than compare one phone, it's more illuminating to watch the gap evolve. I've picked four chips, each representing a distinct architectural era of Apple silicon — and skipped everything before 2013, because a fair FP64 comparison requires a 64-bit machine, and Apple's chips only went 64-bit with the A7.
| Era | Chip | Year | What it represents |
|---|---|---|---|
| 64-bit matures | A9 | 2015 | Dual-core ARMv8, the last pre-heterogeneous |
| design | |||
| Heterogeneous multicore | A11 Bionic | 2017 | The 2+4 big.LITTLE template, |
| first Neural Engine | |||
| The Bionic plateau | A16 Bionic | 2022 | Peak refinement of a |
| five-year-stable design | |||
| Modern 3nm / ARMv9 | A19 Pro | 2025 | Current flagship, iPhone 17 Pro Max |
The honest comparison: matched precision
Here's the row that matters most — floating-point throughput at the same precision the Cray used, FP64, computed from each chip's CPU vector units.
| System | FP64 peak | vs. X-MP/48 |
|---|---|---|
| Cray X-MP/48 | 800 MFLOPS | 1× |
| A9 (2015) | ~44 GFLOPS | ~56× |
| A11 (2017) | ~103 GFLOPS | ~128× |
| A16 (2022) | ~175 GFLOPS | ~219× |
| A19 Pro (2025) | ~220 GFLOPS | ~274× |
Two things jump out. First, even the decade-old A9 already beat the entire four-processor supercomputer by roughly 56× on matched precision — a phone chip from 2015, in your pocket, comfortably ahead. Second, the current flagship's lead is about 274×. Substantial, absolutely. But it is not the "thousands of times faster" you usually hear.
Where these numbers come from. Apple doesn't publish FP64 throughput, so these are derived from the microarchitecture: cores × clock × NEON pipelines × 2 FP64 lanes per 128-bit pipe × 2 for fused multiply-add. The pipeline counts are the load-bearing assumption. For the modern Firestorm-lineage cores (A16, A19 Pro) they're well established through reverse engineering — four SIMD units on the performance cores, two on the efficiency cores. For the older A9 and A11, pipe counts are less firmly documented, so treat those two figures as good estimates with perhaps ±1 pipe of uncertainty rather than precise values. The A19 Pro's ~220 GFLOPS assumes all NEON pipes issue an FMA every cycle at max clock — a true ceiling for FMA-bound code, not a sustained figure.
The row the headlines actually use
So where does "thousands of times faster" come from? From quietly changing precision. The A19 Pro's GPU delivers roughly 2.5 TFLOPS at FP32 and around 5 TFLOPS at FP16 — and its Neural Engine, measured in INT8 operations, reaches tens of TOPS. Line up the GPU's FP32 peak against the Cray's FP64 peak and you get about 3,100×. Use FP16 and it's over 6,000×.
But this compares different things. The Cray's 800 MFLOPS was double precision, the demanding format that scientific and engineering codes actually require. Most of a phone's spectacular throughput lives in lower-precision formats built for graphics and machine learning — formats the X-MP didn't even implement. It's a real capability the phone has and the Cray lacked. It just isn't the same benchmark, and quoting it as though it were flatters the phone by a factor of ten or more.
Memory bandwidth: where the Cray nearly holds the line
Now for a more tempered result. Memory bandwidth was the X-MP's crown jewel — so how does it fare?
| System | Memory bandwidth | vs. X-MP/48 |
|---|---|---|
| Cray X-MP/48 | ~10 GB/s (aggregate) | 1× |
| A9 (2015) | 25.6 GB/s | 2.5× |
| A11 (2017) | 34.1 GB/s | 3.4× |
| A16 (2022) | 51.2 GB/s | 5.1× |
| A19 Pro (2025) | 76.8 GB/s | 7.6× |
Read that last column again. On the single axis where the Cray was a world-beater, the 2025 flagship iPhone leads by less than eightfold — versus 274× on matched compute and thousands-fold on mismatched-precision compute. A phone that outruns the supercomputer by 274× in double-precision arithmetic can feed its processors with only 7.6× the memory bandwidth. The Cray's designers spent their transistor budget on exactly the right thing, and it still shows forty years later.
There's even a latency footnote that goes the wrong way for the phone: the Cray's static main memory answered in about 38 nanoseconds, while modern LPDDR5X DRAM latency is often over 100 ns. On raw latency, the 1984 machine wins.
The storage inversion
One last comparison, because it's too good to leave out. The X-MP offered an optional "SSD" — but in 1984 that meant a Solid-state Storage Device built from volatile MOS RAM, used as a blisteringly fast disk. It moved up to 2,000 MB/s with 25-microsecond access.
The iPhone 17 Pro's NAND flash — nonvolatile, up to 2 TB, sipping milliwatts — delivers measured sequential reads of a little over 2,000 MB/s. Essentially the same sequential bandwidth. The Cray achieved it with a volatile, power-hungry, refrigerator-adjacent bank of RAM roughly a thousand times less dense; the phone does it with nonvolatile storage you'd never think about. Same number, utterly different engineering problem — four decades of progress hiding inside an unchanged headline figure.
The gap that really is astronomical: power
Everything so far has been about capability, where the honest gaps turned out modest. Now flip to efficiency, and the story inverts completely.
The X-MP/48 did not run on a battery. Its mainframe drew on the order of 345 kilowatts, and that figure is only the computer itself — the ECL circuitry ran so hot that Cray's designers spent as much effort on refrigeration as on the processor, and the Freon-and-chilled-water cooling plant roughly doubled the site total to something near half a megawatt. At NCSA we didn't think of it as a computer you plugged in; it was a computer with its own electrical and cooling infrastructure, sitting in a purpose-built room.
The iPhone 17 Pro Max runs on a 19-watt-hour battery. Under a sustained compute load the whole device draws around 10 watts. Asleep on a nightstand, it sips something like 50 milliwatts.
| Cray X-MP/48 | iPhone 17 Pro Max | |
|---|---|---|
| Power, working | ~500 kW (with cooling) | ~10 W |
| Power, idle | ~500 kW (it doesn't rest) | ~50 mW |
| Power source | dedicated substation feed | pocket battery |
That's roughly a 50,000-fold difference in operating power — and unlike the Cray, the phone can throttle to milliwatts the instant it's idle. The supercomputer had no such gear; it burned its half-megawatt whether it was solving equations or waiting for the next job.
Now divide power by work to get the metric that matters: energy per calculation. Running real double-precision workloads, the X-MP/48 spent on the order of a millijoule per FP64 operation. The iPhone, doing the same double-precision work, spends a fraction of a nanojoule — well under one ten-thousandth as much energy per calculation. Put the two together and the phone is something like ten million times more energy-efficient at double-precision arithmetic.
The honest four orders of magnitude. This is where the "thousands of times better" instinct is actually right — it was just pointed at the wrong number. The phone isn't thousands of times faster; matched precision to precision it's a couple hundredfold. But it is millions of times more efficient. Four-plus decades of progress didn't mostly go into raw speed. It went into doing each calculation for almost no energy at all — which is what let the supercomputer leave the refrigerated room and move into your pocket.
What the comparison actually shows
The point isn't that the iPhone is unimpressive. A pocket device that beats the world's fastest 1984 supercomputer by 274× in double precision — while doing it on ten watts instead of half a megawatt, with no refrigerated room — is a staggering achievement. But notice where the achievement actually lives. Not in raw speed, which improved a couple hundredfold. In efficiency, which improved ten-million-fold. Those two numbers are the whole story.
But raw capability is a more modest story than the marketing math implies, and an honest one is more interesting than "a phone is 5,000× a Cray." Match precision to precision and the gap is a couple hundredfold. Look at memory bandwidth — the thing the Cray was actually designed around — and the machines are within one order of magnitude. The supercomputer earned its reputation on fundamentals that still hold up. It was never really about the FLOPS. It was about keeping those FLOPS fed.
One more note: portability. The XMP/48 weighted in at a svelte 5.12 tons. Even if the power requirements didn't repvent portability, the sheer size definitely did.
Figures for the X-MP/48 are drawn from Cray's own 1985 documentation and contemporary technical sources; iPhone figures combine Apple's published specifications with microarchitecture analysis and measured benchmarks. FP64 throughput for Apple chips is derived rather than published, with the caveats noted above.