Lecture 05 - Multicycle Operations
Course: Computer Architectures
Professor: E. Sanchez
Source: Multicycle operations, ASE 2026/27, 42 slides
Numbering: Lecture 05 follows the naming agreed in the earlier chat; the PDF itself has no lecture number.
What the professor covered
Multicycle execution (slides 2-10)
Complex FP operations need several clock cycles to avoid a slow clock or excessively complex single-cycle units. EX becomes several functional units with different execution times.
The example has a one-stage integer ALU, a four-stage FP adder (A1...A4), a seven-stage FP/integer multiplier (M1...M7), and an unpipelined FP/integer divider (D). These paths converge on MEM and WB.
The professor distinguishes result-use timing from the spacing between starts:
| Functional unit | Latency | Initiation interval |
|---|---|---|
| Integer ALU | 0 | 1 |
| Data Memory | 1 | 1 |
| FP add | 3 | 1 |
| FP/integer multiply | 6 | 1 |
| FP/integer divide | 24 | 25 |
Use the lecture’s latency convention; 0 does not mean instantaneous execution. See Latency and Initiation Interval and Multicycle Operations.
Hazards and issue checks (slides 11-20)
- Structural: divider occupancy and simultaneous register writes. Slide 12 shows multiply, add, and load reaching WB in cycle 11. Extra write ports are normally too expensive; stall at ID or before MEM/WB.
- RAW: the
fld -> fmul.d -> fadd.d -> fsdchain illustrates long waits for results. - WAW: an older
fadd.dand youngerfldboth writef2; varying execution times threaten write order.
If detection is performed in ID, check resource availability, source registers against pending destinations and their availability times, and the new destination against destinations in A1...A4, D, and M1...M7. Stall when unsafe.
See Structural Hazards in Multicycle Pipelines, RAW and WAW Hazards, and Hazard Detection in ID.
MIPS R4000 (slides 21-38)
The lecture introduces the 64-bit R4000 (1991) as an eight-stage superpipeline:
IF -> IS -> RF -> EX -> DF -> DS -> TC -> WB
Pipelined instruction/data memory accesses support one new instruction per cycle, but require more forwarding. The professor gives a load delay slot of 2 cycles and branch delay slot of 3 cycles. Load data is forwarded from the end of DS; branch evaluation occurs in EX.
The R4000 FP section uses stage types A, D, E, M, N, R, S, U and a separate operation-timing table. Slide 38 reduces latency by 1 when the consuming instruction is a store.
See MIPS R4000 Pipeline, Load and Branch Delay Slots, and MIPS R4000 FP Pipeline.
Performance (slides 39-42)
In the shown SPEC92 results, branch delays dominate integer-program stall costs, while FP result stalls dominate FP-program stall costs. The professor reports total CPI between 1.2 and 2.8, depending on the program.
See Pipeline CPI and Stall Costs.
Questions to check understanding
- Why can a multiplier have latency 6 but initiation interval 1?
- How can instructions issued in order reach WB out of order?
- What distinguishes two writes to one register from two writes needing one port?
- Why does forwarding still leave a two-cycle load delay in the R4000 example?
Source
E. Sanchez, Multicycle operations, Politecnico di Torino, ASE 2026/27, slides 1-42. PDF page numbers match slide numbers.