Benchmarking Alpha: SPEC Results in Context
For most of the 1990s, if you wanted the fastest floating-point processor you could buy, the answer printed at the top of the SPEC tables was usually an Alpha. This piece puts those results in context: what the benchmarks measured, why Alpha did so well, how much of the credit belonged to compilers, and what to be careful about when quoting old numbers.
What SPEC measured
The Standard Performance Evaluation Corporation's CPU suites were the common currency of workstation and server marketing. Two generations matter for Alpha:
- CPU95 (SPECint95 and SPECfp95), used from the mid-1990s until about 2000. Scores were ratios against a reference machine, a Sun SPARCstation 10/40, which scored 1.
- CPU2000 (SPECint2000 and SPECfp2000), released at the end of 1999, with a Sun Ultra 5/10 at 300 MHz as the reference, scaled so that it scored 100.
Each suite was a collection of real programs rather than synthetic kernels: compilers, compression, chess, circuit simulation and interpreters on the integer side; fluid dynamics, weather modelling, quantum chemistry and image processing on the floating-point side. The composite score was a geometric mean of the per-benchmark ratios.
Each result also came in two flavours. Base required the same compiler flags for every program in the suite. Peak allowed per-program tuning. The gap between the two is a useful measure of how much a result depends on compiler heroics.
How Alpha led floating point
From the 21064 onward, Alpha was designed for high clock speeds and wide data paths, and it led the frequency race for much of the decade. The 21164 paired that with a well-balanced pipeline and on-chip secondary cache, and the 21264 added out-of-order execution, two floating-point pipelines and a far more capable memory system.
Approximate, hedged figures give a sense of the trajectory. These are orders of magnitude drawn from our reading of the period, and exact published results for specific systems should be checked against SPEC's own archive:
| Processor (approx. clock) | Era | SPECint (approx.) | SPECfp (approx.) |
|---|---|---|---|
| 21164 (~500 MHz) | CPU95, c. 1996-97 | Mid-teens | Around 20 |
| 21264 (~500-575 MHz) | CPU95, c. 1998 | Roughly 25-30 | Roughly 40-55 |
| 21264A (~667-833 MHz) | CPU2000, c. 2000 | Roughly 400-550 | Roughly 500-650 |
| 21264C / EV7 (~1-1.25 GHz) | CPU2000, c. 2002-04 | Roughly 700-900 | Roughly 1,100-1,500+ |
The pattern matters more than any individual number. Alpha's integer scores were competitive and often leading, but its floating-point scores tended to lead by a wider margin, particularly on memory-hungry programs. That was the part of the market, technical and scientific computing, where Alpha found its most loyal customers and where API's boards and clusters were sold.
Compilers: Compaq C and Fortran versus GCC
A SPEC result is a measurement of a processor, a memory system and a compiler together. On Alpha, the compiler contribution was unusually large.
Digital, later Compaq, invested heavily in its own compilers: the GEM optimising back end shared by DEC C, Compaq C, C++ and Fortran. These compilers understood the 21164 and 21264 pipelines in detail, scheduled instructions for them, performed aggressive loop transformations and inlining, and supported profile-directed feedback. A post-link optimiser, Spike, could then rearrange the final binary for better instruction-cache behaviour. Tuned math libraries, notably CXML, supplied hand-optimised routines.
GCC in the late 1990s produced correct Alpha code but, for floating-point work especially, often noticeably slower code than Compaq's tools. Gaps of tens of percent were commonly reported on numeric workloads, and occasionally more. Recognising how much this mattered to the Linux market, Compaq released versions of its C and Fortran compilers and math library for Alpha Linux, and cluster integrators generally shipped them.
Memory bandwidth and the STREAM benchmark
Many scientific codes are limited less by arithmetic than by how quickly the processor can move data between memory and registers. The STREAM benchmark, a small program measuring sustainable bandwidth on simple vector kernels (copy, scale, add and triad), became the standard way to quantify this.
Here Alpha systems looked especially strong. The 21264's EV6 bus, the same point-to-point bus design licensed to AMD for the Athlon (see our EV6 bus article), paired with wide, interleaved memory on Tsunami-class chipsets, delivered sustained bandwidth that was, roughly, a gigabyte per second or more per system in the ES40 era, well ahead of typical PC-class platforms of the time. The EV7 generation later integrated memory controllers on the chip itself and reported per-processor STREAM figures several times higher again.
That bandwidth goes a long way to explaining the floating-point lead: several SPECfp programs are effectively memory bandwidth tests with some arithmetic attached. It also explains why memory population mattered so much, as covered in our memory configuration guide.
Methodology caveats
Old benchmark numbers are easy to quote and easy to misuse. Before drawing conclusions, keep the following in mind:
- Base versus peak. Always know which one you are comparing. Mixing a base result from one vendor with a peak result from another was a favourite marketing trick.
- Suites are not comparable across generations. A CPU95 score and a CPU2000 score use different programs and reference machines; there is no reliable conversion factor.
- System, not chip. Results depend on cache size, memory configuration, chipset and disk. Two machines with the same processor could differ substantially.
- Compiler special-casing. Several vendors, not only Digital and Compaq, were accused of optimisations that targeted particular SPEC programs more than general code. Individual benchmarks, such as the art component of CPU2000, saw dramatic jumps that owed more to compilers than hardware.
- Single-thread focus. The speed metrics measured one copy of each program. Rate metrics measured throughput with many copies, and are the better guide for cluster nodes.
- Price and power were absent. SPEC said nothing about cost or watts, which is where Alpha systems were most vulnerable against cheaper x86 machines.
Reading the results in 2026
With all those caveats, the broad story holds up. Through the late 1990s Alpha delivered the highest floating-point performance in the industry more often than any rival, on the strength of clock speed, a balanced memory system and excellent compilers. By the early 2000s, x86 processors with rapidly improving memory systems were closing the gap at a fraction of the price, a shift our EV7 and EV8 roadmap article discusses. Alpha's final SPEC results were still excellent; they simply stopped being decisive.
Further reading on this archive
- 21264A processor page: the chip behind many of API's CPU2000-era results.
- The Alpha Linux Beowulf cluster era: where benchmark performance met real scientific workloads.
- All archive posts: browse the rest of the blog.