Pipeline CPI and Stall Costs
CPI measures the average number of clock cycles per completed instruction. Pipeline stalls increase it.
What the lecture shows
Slides 39-42 show stacked pipeline-CPI bars for SPEC92 benchmarks, with these components:
- Base
- Load stalls
- Branch stalls
- FP result stalls
- FP structural stalls
The professor highlights branch delays as the main stall contribution for integer programs and FP result stalls as the main stall contribution for FP programs. The reported total CPI range is 1.2 to 2.8, depending on the program.
These are observations about the workloads and pipeline shown, not a universal CPI range or bottleneck for every processor.
Reusable explanation: accounting for stalls
For a scalar pipeline whose ideal steady-state CPI is 1, a useful model is:
If 100 completed instructions require 100 base cycles and 40 extra stall cycles, the CPI is 1.4, ignoring fill/drain effects in this simple example. Count a stall cycle once even if several hazards overlap during it.
Execution time also depends on the clock period:
A deeper pipeline can shorten the clock period while increasing hazard penalties. CPI alone therefore does not establish which processor is faster.
Why the bottlenecks differ
Integer workloads in the shown results lose cycles waiting for branch outcomes. FP workloads lose many cycles waiting for long-latency results used by dependent instructions. A unit accepting new independent work quickly does not eliminate dependent-result stalls.
Related: Load and Branch Delay Slots, RAW and WAW Hazards, Latency and Initiation Interval, MIPS R4000 FP Pipeline.
Source
E. Sanchez, Multicycle operations, Politecnico di Torino, ASE 2026/27, slides 39-42. PDF page numbers match slide numbers.
Lecture context: Lecture 05 - Multicycle Operations. Sections marked as clarification or reusable explanation add study guidance to the slide material.