Alpha Linux and the Beowulf Era of Cluster Computing
In the late 1990s, if you wanted the most floating-point work per dollar out of a rack of commodity machines running Linux, the answer was often Alpha. This is a look back at why that was true, how those clusters were wired together and where API fitted into the picture.
From Beowulf to a movement
The original Beowulf cluster, built at NASA Goddard in 1994, was sixteen Intel 486 machines connected by Ethernet and running Linux. Its point was not raw speed but economics: ordinary PC parts, a free operating system and message-passing libraries such as PVM and later MPI could deliver useful parallel performance at a fraction of the cost of a proprietary supercomputer. Within a few years, universities and national laboratories were building their own clusters, and "Beowulf" became shorthand for the whole approach.
The recipe was simple enough to fit on a slide:
- Many identical nodes, each a complete computer with its own memory and disk.
- A private network between them, as fast as the budget allowed.
- Linux, a compiler toolchain and an MPI implementation.
- A head node for logins, job scheduling and shared file systems.
What varied was the choice of node and the choice of network, and those two decisions determined almost everything about how a cluster performed.
Why Alpha won on floating point
Scientific codes of the era, such as computational fluid dynamics, molecular dynamics, weather models and dense linear algebra, were dominated by double-precision floating-point arithmetic and memory bandwidth. On both counts, Alpha was ahead of contemporary x86 parts by a wide margin.
The 21164A could issue two floating-point operations per clock and, at 500 to 600 MHz, posted SPECfp95 results well above those of Pentium II systems at similar prices. The 21264 and 21264A widened the gap with out-of-order execution, a large on-chip data cache, a big off-chip B-cache and a much faster system bus. x86 floating point at the time still ran through the x87 stack, and SSE2 double-precision vector support did not arrive until the Pentium 4 in late 2000.
| Node type (c. 1999) | Clock | Peak DP GFLOPS | Relative node cost |
|---|---|---|---|
| Pentium II / III | 400–550 MHz | ~0.4–0.55 | Low |
| Alpha 21164A | 533–600 MHz | ~1.1–1.2 | Moderate |
| Alpha 21264 / 21264A | 500–750 MHz | ~1.0–1.5 | Moderate to high |
Peak numbers flatter everyone, but sustained results told a similar story: on bandwidth-hungry codes an Alpha node was frequently worth two or more x86 nodes, which also halved the number of network ports, cables and failure points. For many labs, that tipped the balance even when each Alpha node cost more.
Landmark Alpha clusters
A handful of systems made the case publicly. Los Alamos's Avalon, built in 1998 from roughly 140 Alpha 21164A nodes on Fast Ethernet, entered the TOP500 list somewhere in the low hundreds, which was remarkable for a machine assembled from commodity parts on a small budget. Sandia's CPlant project scaled Alpha Linux clusters with Myrinet into the high hundreds and eventually well over a thousand nodes, and several of its incarnations ranked comfortably inside the TOP500.
At the very top of the list, Compaq's AlphaServer SC machines, which ran Tru64 rather than Linux and used the Quadrics interconnect, reached the top few positions around 2001 and 2002. They were not Beowulf clusters in the strict sense, but they showed that clustered Alpha nodes could compete with anything in the world at scale.
The interconnect question
Fast nodes are wasted if they spend their time waiting for messages. Switched Fast Ethernet was cheap, but its latency, often 60 to 100 microseconds through the TCP/IP stack, crippled tightly coupled codes. Serious clusters looked elsewhere.
Myrinet
Myricom's Myrinet was the default high-performance choice for Alpha clusters. It offered gigabit-class links, cut-through switches and user-level messaging software that bypassed the kernel, bringing MPI latencies down to the order of ten microseconds. Its downside was cost: a Myrinet adapter and switch port could approach the price of a modest node.
SCI and WulfKit
The Scalable Coherent Interface took a different route. Dolphin's SCI adapters connected nodes in rings or two- and three-dimensional tori without a central switch, and Scali's MPI and management software turned the hardware into a turnkey package called WulfKit. In August 2000 API announced support for Dolphin/Scali WulfKit on its Alpha platforms, offering single-digit microsecond latencies and a cluster that scaled by adding cables rather than switch capacity.
A typical job launch on such a system looked no different from any MPI cluster:
$ mpicc -O3 -o solver solver.c -lm
$ mpirun -np 32 -machinefile nodes.txt ./solver input.datAPI's role: boards, partners and integrators
Alpha Processor, Inc. was well placed to sell into this market. Its UP1000 and UP2000 boards brought Alpha to ATX form factors and commodity cases, and the CS20, announced in November 2000, packed two 21264 processors into a 1U enclosure designed specifically for racks of cluster nodes. What API lacked was the integration business: the people who would design, cable, burn in and support a hundred-node machine.
It solved that through partners. In August 2000 API announced work with Linux NetworX, a Utah-based cluster integrator, to build Alpha Linux clusters with management tools and support. In January 2001 came an alliance with Cray, which gave API's Alpha Linux cluster offerings a supercomputing brand and sales channel, and gave Cray a commodity cluster line alongside its vector and massively parallel machines.
Why the Alpha cluster era ended
Several forces converged between 2001 and 2003:
- Compaq's 2001 decision to move its high-end roadmap to Itanium removed long-term confidence in Alpha.
- The Pentium 4 with SSE2, and then the Athlon 64 and Opteron, closed most of the floating-point gap at much lower prices.
- Opteron in particular offered 64-bit addressing, integrated memory controllers and HyperTransport, directly addressing the reasons labs had chosen Alpha.
- Cluster software had matured to the point where the node's instruction set mattered less than its price-performance.
By the middle of the decade, x86 clusters dominated the TOP500, and they have never really given that position up. The model that Alpha Linux clusters helped prove, commodity nodes with fast interconnects and free software, became the model for nearly every supercomputer that followed.
Looking back from 2026
It is easy to forget how contrarian the Beowulf idea once was. Alpha's contribution was to show that commodity clusters did not have to mean commodity performance: a well-built Alpha Linux cluster could hold its own against proprietary machines costing many times more. If you run a small Alpha cluster today, whether on real UP2000 or CS20 hardware or under emulation, you are reconstructing one of the more consequential experiments in computing history.
Further reading on this archive
- Running Linux on Alpha hardware today covers distributions and kernels for building nodes.
- Benchmarking Alpha: SPEC in context puts the floating-point numbers above in perspective.
- The CS20 and dense 1U computing looks at API's purpose-built cluster node.