Lecture 04 - Pipelining
Course: Computer Architectures
Source: Pipelining lecture slides
Topics
Lecture map
The lecture defines pipelining as overlapping the execution of multiple instructions, with different stages processing different instructions in parallel.
The five stages used in the lecture are:
IF Instruction Fetch
ID Instruction Decode / Register Fetch
EX Execute / Effective Address
MEM Memory Access / Branch Completion
WB Write BackFor the detailed study note, use Pipelining.
Performance points emphasized in the lecture
- Pipelining increases throughput rather than making a single instruction faster.
- Pipeline stages are synchronized.
- The clock period is constrained by the slowest stage.
- In the ideal case, a filled pipeline approaches one completed instruction per cycle.
- Pipeline depth is limited by stage balancing and pipeline overhead.
Pipeline registers
The pipelined datapath uses registers between stages:
IF/ID
ID/EX
EX/MEM
MEM/WBThese keep the state belonging to different in-flight instructions separated.
Three hazard classes
The lecture classifies hazards as:
Structural → resource conflict
Data → dependency on an earlier result
Control → branch/jump changes the PCStructural hazards
Example from the lecture: a single-port memory may be needed by both instruction fetch and a load/store in the same clock cycle.
Possible response:
stallor improve/add hardware.
Data hazards
Example dependency:
ADD R1, R2, R3
SUB R4, R1, R5The second instruction may read R1 before the first has written the new value.
Forwarding
The lecture solves many data hazards by forwarding results from pipeline registers directly to functional-unit inputs.
Conceptually:
previous ALU result ─────→ current ALU inputLoad-use hazard
The lecture stresses that not all hazards are solved by forwarding.
Example:
LD R1, 0(R2)
SUB R4, R1, R5The loaded value is not available early enough for the immediately following instruction.
A stall is therefore required.
What happens during that stall
The lecture’s control mechanism can:
- force a nop into ID/EX
- keep IF/ID unchanged
- freeze the PCIn RISC-V:
nopcorresponds to:
addi x0, x0, 0Control hazards
The lecture defines control hazards as hazards due to branches or other instructions that change the PC after later instructions may already have been fetched.
In the basic implementation presented, the branch outcome/target affects the PC at the end of EX, two clock cycles after the branch’s IF stage.
This means a taken branch can leave two younger instructions on the wrong path.
Taken vs not taken
For a conditional branch:
taken → PC is changed to the target
not taken → sequential execution continuesIn the lecture’s basic example, a taken branch loses two cycles because the two younger sequential instructions must be discarded.
Branch-management techniques
The lecture presents:
1. Freeze the pipeline
2. Predict untaken
3. Predict taken
4. Delayed branchFreeze
Wait until the decision is known.
Predict untaken
Continue along the sequential path.
If the branch later turns out to be taken, undo/flush the incorrect work.
Predict taken
If the target address is known early enough, begin fetching from the target.
Delayed branch
Use the instruction slot after a branch for work that is valid regardless of the outcome.
The compiler is responsible for finding a suitable instruction.
Professor-specific points worth remembering
- The example machine uses a classic five-stage pipeline.
- Pipeline hazards are divided into structural, data, and control.
- A stall introduces a bubble.
- Forwarding solves many, but not all, data hazards.
- An immediate load-use dependency requires a stall.
- The lecture’s basic branch implementation can lose two cycles for a taken branch.
- Branch handling techniques compared: freeze, predict untaken, predict taken, delayed branch.
Study links
- Pipelining > 4. Pipeline hazards
- Pipelining > 8. Forwarding
- Pipelining > 9. Load-use hazard
- Pipelining > 11. Control hazards
- Pipelining > 17. Predict untaken
- Pipelining > 18. Predict taken
- Pipelining > 20. Delayed branch
Questions to review
- Why does pipelining improve throughput rather than instruction latency?
- Why is the clock period determined by the slowest stage?
- What is a bubble?
- Which hazard can forwarding solve?
- Why does a load-use dependency still require a stall?
- Why can a taken branch waste instructions already in the pipeline?
- What is the difference between flushing and stalling?