Lecture 04 - Pipelining

Course: Computer Architectures
Source: Pipelining lecture slides

Topics

Lecture map

The lecture defines pipelining as overlapping the execution of multiple instructions, with different stages processing different instructions in parallel.

The five stages used in the lecture are:

IF  Instruction Fetch
ID  Instruction Decode / Register Fetch
EX  Execute / Effective Address
MEM Memory Access / Branch Completion
WB  Write Back

For the detailed study note, use Pipelining.

Performance points emphasized in the lecture

  • Pipelining increases throughput rather than making a single instruction faster.
  • Pipeline stages are synchronized.
  • The clock period is constrained by the slowest stage.
  • In the ideal case, a filled pipeline approaches one completed instruction per cycle.
  • Pipeline depth is limited by stage balancing and pipeline overhead.

Pipeline registers

The pipelined datapath uses registers between stages:

IF/ID
ID/EX
EX/MEM
MEM/WB

These keep the state belonging to different in-flight instructions separated.

Three hazard classes

The lecture classifies hazards as:

Structural → resource conflict
Data       → dependency on an earlier result
Control    → branch/jump changes the PC

Structural hazards

Example from the lecture: a single-port memory may be needed by both instruction fetch and a load/store in the same clock cycle.

Possible response:

stall

or improve/add hardware.

Data hazards

Example dependency:

ADD R1, R2, R3
SUB R4, R1, R5

The second instruction may read R1 before the first has written the new value.

Forwarding

The lecture solves many data hazards by forwarding results from pipeline registers directly to functional-unit inputs.

Conceptually:

previous ALU result ─────→ current ALU input

Load-use hazard

The lecture stresses that not all hazards are solved by forwarding.

Example:

LD  R1, 0(R2)
SUB R4, R1, R5

The loaded value is not available early enough for the immediately following instruction.

A stall is therefore required.

What happens during that stall

The lecture’s control mechanism can:

- force a nop into ID/EX
- keep IF/ID unchanged
- freeze the PC

In RISC-V:

nop

corresponds to:

addi x0, x0, 0

Control hazards

The lecture defines control hazards as hazards due to branches or other instructions that change the PC after later instructions may already have been fetched.

In the basic implementation presented, the branch outcome/target affects the PC at the end of EX, two clock cycles after the branch’s IF stage.

This means a taken branch can leave two younger instructions on the wrong path.

Taken vs not taken

For a conditional branch:

taken     → PC is changed to the target
not taken → sequential execution continues

In the lecture’s basic example, a taken branch loses two cycles because the two younger sequential instructions must be discarded.

Branch-management techniques

The lecture presents:

1. Freeze the pipeline
2. Predict untaken
3. Predict taken
4. Delayed branch

Freeze

Wait until the decision is known.

Predict untaken

Continue along the sequential path.

If the branch later turns out to be taken, undo/flush the incorrect work.

Predict taken

If the target address is known early enough, begin fetching from the target.

Delayed branch

Use the instruction slot after a branch for work that is valid regardless of the outcome.

The compiler is responsible for finding a suitable instruction.

Professor-specific points worth remembering

  • The example machine uses a classic five-stage pipeline.
  • Pipeline hazards are divided into structural, data, and control.
  • A stall introduces a bubble.
  • Forwarding solves many, but not all, data hazards.
  • An immediate load-use dependency requires a stall.
  • The lecture’s basic branch implementation can lose two cycles for a taken branch.
  • Branch handling techniques compared: freeze, predict untaken, predict taken, delayed branch.

Questions to review

  • Why does pipelining improve throughput rather than instruction latency?
  • Why is the clock period determined by the slowest stage?
  • What is a bubble?
  • Which hazard can forwarding solve?
  • Why does a load-use dependency still require a stall?
  • Why can a taken branch waste instructions already in the pipeline?
  • What is the difference between flushing and stalling?