The CS20 and the Birth of Dense 1U Computing

In November 2000, API announced the CS20: two Alpha 21264 processors in a chassis a single rack unit tall. Today 1U dual-socket servers are the default building block of every data centre, but at the time fitting that much floating-point performance into 1.75 inches of height was an engineering statement.

The machine

The CS20, announced in this November 2000 press release, packaged two 21264-family processors, in the 833 MHz range at launch, into a 1U rack-mount enclosure aimed squarely at clusters and technical computing. It drew on the same chipset lineage as API's UP2000 dual-processor board, reworked onto a board shape that could sit flat in a very shallow box.

Exact specifications varied with configuration, and the numbers below should be treated as approximate, but a typical CS20 offered:

  • Two 21264-class Alpha CPUs, each with a large off-chip backside cache.
  • Registered ECC SDRAM, up to a few gigabytes.
  • A small number of PCI expansion slots via riser, enough for a cluster interconnect card.
  • Integrated Ethernet, room for one or two drives, and SRM console firmware for Linux and Tru64.

That list is unremarkable now. What made it hard was the 21264 itself: a large, hot, high-frequency chip designed originally for deskside workstations and tall servers with generous airflow.

Thermal engineering in 1.75 inches

A tower or 3U server can use tall heatsinks and large, slow fans. In 1U, heatsink height is capped at well under 40 mm once you allow for the board and lid, and the only way to remove heat is to push a lot of air quickly through a narrow channel. The CS20's layout reflects that constraint:

  1. Front-to-back airflow. Cool air enters at the front, passes over drives and then processors, and exits at the rear. Every component is arranged so it does not shadow another from the airflow.
  2. Low-profile heatsinks with dense fins. Since the sink cannot be tall, it must be wide and finely finned, which in turn needs high static pressure from the fans.
  3. Banks of small, fast fans. Multiple fans in a row across the chassis provide redundancy and pressure. The cost is noise, which is why 1U servers earned their reputation for sounding like jet engines.
  4. Ducting. Plastic or sheet-metal baffles force air through the heatsinks instead of around them, a detail that is easy to lose when a machine is serviced carelessly.

Power, and why it became the real limit

Each 21264 dissipated on the order of 70 to 90 watts depending on speed grade, and memory, chipset, drives and conversion losses roughly doubled the box's draw. A fully configured CS20 plausibly consumed somewhere in the region of 250 to 400 watts at load. Per box, that is fine. Per rack, it changes the conversation.

In 2000 a typical enterprise machine room was provisioned for perhaps 2 to 5 kilowatts per rack, because racks were filled with a few tall servers, storage and network gear. Dense 1U compute turned that assumption upside down.

Rack density math

The table below uses deliberately round, hedged figures to show what happened as form factors shrank. It assumes a standard 42U rack, and reserves some space for switches and cable management.

Form factorServers per rack (approx.)CPUs per rackPower per server (est.)Rack load (est.)
4U dual-CPU9-1018-20~400 W~4 kW
2U dual-CPU18-2036-40~350 W~7 kW
1U dual-CPU (CS20 class)36-4072-80~300 W~11-12 kW
Modern 1U dual-socket (2026, typical)36-4072-80 sockets, thousands of cores~700-1,000+ W~25-40 kW

The punchline is that the CS20 class of machine could put roughly four times the processors of a 4U configuration into the same floor tile, and in doing so tripled or quadrupled the heat that tile had to reject. Many sites that bought dense clusters in 2000 and 2001 found they could fill only half a rack before running out of power circuits or cooling capacity, and ended up with half-empty racks spread across more floor space. That tension between density and facility capacity has never gone away; it has simply moved up by an order of magnitude.

Cluster deployments

The CS20 was built for clustering, and it landed in a market that API had been cultivating for a year. Integrators such as Linux NetworX, whose Alpha cluster work is recorded on this archive, assembled racks of Alpha nodes running Linux for scientific and engineering customers. High-performance interconnects mattered as much as the nodes, and the Dolphin and Scali WulfKit announcement from August 2000 shows the kind of SCI-based networking that was paired with Alpha systems to cut message latency well below what Ethernet could manage.

The typical CS20 deployment looked like this:

  • Dozens to a few hundred nodes, each running Linux from local disk or a network boot.
  • A dedicated low-latency interconnect for MPI traffic, plus Ethernet for management and file access.
  • Workloads dominated by floating-point simulation: computational fluid dynamics, seismic processing, structural analysis and physics codes where the 21264's floating-point and memory performance paid off.

Our Beowulf cluster era article covers the software side: the MPI stacks, schedulers and compilers that made these machines usable.

Legacy for modern 1U and blade servers

The CS20 was not the only early dense server, and it would be an overreach to credit it with inventing the category. x86 vendors were shipping 1U machines in the same period, and dedicated blade systems appeared around 2001. What the CS20 demonstrated was that a genuinely high-end, hot, RISC processor could be packaged this way, and that the resulting cluster node was a sensible unit of purchase for technical computing.

Several of its design problems became industry-standard concerns:

  • Airflow as architecture. Front-to-back cooling and hot-aisle/cold-aisle layouts became standard practice because dense servers demanded them.
  • Facility-limited density. Watts per rack, not units per rack, became the planning metric, a shift that ultimately led to today's liquid-cooled racks.
  • Shared infrastructure. Blade chassis later pooled fans, power and networking across many nodes, directly addressing the redundancy and cabling overhead of stacks of individual 1U boxes.

Flagged speculation: had Alpha development continued, a HyperTransport-connected successor to the CS20 would likely have pushed further into blade-style shared chassis, given API's parallel work on HyperTransport bridging. The roadmap ended before that could be tested.

Further reading on this archive

← Back to the archive