Lab 01 - gem5 and ASE Studio
Course: Computer Architectures
Laboratory: 0x01
Deadline: 06/10/2026
Status: Programs and measurements completed; PDF and submission pending
Submission
Expected submission:
SURNAME_NAME_MATRICOLA_lab_01.zipThe archive must include:
program_1program_2lab_01.pdf
Important
Complete the laboratory document and export it as
lab_01.pdfbefore creating the final submission archive.
Objective
Use ASE Studio and gem5 to:
- write and run RISC-V assembly programs;
- configure the simulated CPU;
- visualize pipeline execution;
- collect performance measurements;
- compute CPI, IPC, and execution times;
- compare workload scenarios.
Related concepts
Environment Setup
1. Open a terminal
Open a terminal on the Linux-based machine.
2. Download the simulation environment
ASE Studio is included as a Git submodule of the ase_riscv_gem5_sim repository.
Do not clone the ASE Studio repository separately.
Move to the directory where you want to install the environment and run:
git clone --branch main --recurse-submodules https://github.com/cad-polito-it/ase_riscv_gem5_sim.git
cd ase_riscv_gem5_simThe --recurse-submodules option also downloads ASE Studio inside the parent repository.
3. Install the required tools
Labinf machine
gem5 and the RISC-V toolchain are already installed. Install only ASE Studio:
./utils/installation.sh labinfLocal virtual machine
All required tools are already installed.
Local machine
Install all required tools:
./utils/installation.sh all-aseThe complete installation may take several minutes depending on the machine and Internet connection.
4. Launch ASE Studio
Launch ASE Studio from the desktop shortcut, if available, or from the terminal:
./ase-studio.sh5. Create a new program
In ASE Studio:
- Click New.
- Enter the project name required by the exercise.
- Click Create.
Example:
program_16. Configure the CPU
Configure the CPU according to the configuration specified in the exercise.
CPU settings used for the measurements are recorded below from the saved gem5 run commands.
7. Run the simulation
ASE Studio compiles the program and runs it with gem5 using the selected CPU configuration.
8. Analyze the pipeline
After the simulation finishes:
- Open the pipeline visualization.
- Inspect how instructions move through the pipeline stages.
- Use the visualization to answer the exercise questions.
9. Configure submission settings
Open Settings in ASE Studio.
In Before assignment name, enter:
SURNAME_NAME_MATRICOLAThe After assignment name field does not need to be filled unless explicitly requested.
ASE Studio will use these settings when generating the final submission filename.
10. Prepare the submission
Complete all exercises and export the finished laboratory document as:
lab_01.pdfThen open Prepare submission in ASE Studio.
Set:
Assignment name: lab_01Select these projects:
program_1
program_2Under Additional files, attach:
lab_01.pdfFinally, click Create submission.
Expected result:
SURNAME_NAME_MATRICOLA_lab_01.zipCPU Configuration
Measurements recorded on 2026-10-03. The user confirmed that the exercise CPU settings had already been configured. The following values come from the actual saved gem5 commands, rather than an independent check of the official CPU specification.
| Setting | program_1 | program_2 | calculate_pi | insertion_sort |
|---|---|---|---|---|
| CPU model | MinorCPU | MinorCPU | MinorCPU | MinorCPU |
| Forwarding | on | off | on | off |
| Floating ALU latency (cycles) | 2 | 2 | 2 | 2 |
| Floating multiply latency (cycles) | 7 | 5 | 7 | 7 |
| Floating divide latency (cycles) | 8 | 8 | 8 | 3 |
Common settings: integer ALU/multiply/divide latency 1 cycle; all listed arithmetic units pipelined except floating divide; direct memory; instruction/read/write latency 1 cycle; cache stalls off. The commands specify 32 KiB instruction/data caches, 64-byte cache lines, cache latency 2 and memory latency 30, while using direct-memory mode. The simulated CPU and system clocks are 1 GHz. The exercise’s 1.75 kHz is used afterward to convert measured cycles into the requested execution times; the simulator clock was not changed to 1.75 kHz.
See Lab 01 - Program Sources for the exact measured assembly and saved simulator commands for configuration evidence.
Exercise 1 - Fibonacci
Project name:
program_1Goal
Write and run a RISC-V assembly program that calculates the Fibonacci sequence.
Starting assembly template
# Data section
.section .data
# Place here your program data. In this example
# two vector of floats, a vector of ints and a single int are defined
T0: .word 0x0BADC0DE
# Code section
.section .text
# The _start label signals the entry point of your program
# DO NOT CHANGE ITS NAME. It must be "_start", not "start",
# not "main", not "start_".
# It's "_start" with a leading underscore and
# all lowercase letters
.globl _start
_start:
Main:
# Initialize Fibonacci variables
li x1, 0 # x1 = a = first Fibonacci number (0)
li x2, 1 # x2 = b = second Fibonacci number (1)
li x3, 21 # x3 = count = number of terms to generate
li x4, 0 # x4 = i = loop counter
# Loop to generate and print remaining 20 numbers
addi x4, x4, 1 # i = 1 (start from second iteration)
fib_loop:
beq x4, x3, End # if i == count, exit loop
# Calculate next Fibonacci number
add x5, x1, x2 # x5 = next = a + b
# Update variables for next iteration
mv x1, x2 # a = b (previous second becomes first)
mv x2, x5 # b = next (calculated next becomes second)
# Increment counter and continue loop
addi x4, x4, 1 # i++
j fib_loop # Jump back to loop start
End:
# exit() syscall. This is needed to end the simulation gracefully
li a0, 0
li a7, 93
ecallEquivalent C pseudocode
/*
* Fibonacci sequence generator in C (pseudocode)
*
* Generates the first 21 Fibonacci numbers using iterative approach
*/
int main() {
int i; // Loop counter
int a = 0; // First Fibonacci number
int b = 1; // Second Fibonacci number
int next; // Next Fibonacci number
int count = 21; // Number of terms to generate
// Generate and print remaining numbers
for (i = 1; i < count; i++) {
next = a + b; // Calculate next Fibonacci number
a = b; // Update: previous second becomes first
b = next; // Update: calculated next becomes second
}
return 0;
}Tasks
- Create project
program_1 - Insert the Fibonacci assembly program
- Configure the CPU
- Run the simulation
- Inspect the pipeline
- Record clock cycles
- Record executed instructions
- Compute CPI
- Compute IPC
Results
| Metric | Value |
|---|---|
| Clock cycles | 151 |
| Number of instructions | 126 |
| CPI | 1.1984 |
| IPC | 0.8344 |
Notes / observations
Exercise 2 - Array matching and ordering flags
Project name:
program_2Goal
Write and run a RISC-V assembly program that compares two arrays of signed 8-bit integers and builds a third array containing matching values.
Input arrays
Each array contains 10 signed 8-bit integers.
Example:
v1: .byte 2, 6, -3, 11, 9, 18, -13, 16, 5, 1
v2: .byte 4, 2, -13, 3, 9, 9, 7, 16, 4, 7For each element of v1, check whether it appears in v2 at least once.
Store matching values in v3.
Expected result for the example:
v3: .byte 2, 9, -13, 16Flags
Create three 8-bit unsigned flags.
flag1
flag1 = 1 if v3 is empty
flag1 = 0 otherwiseflag2
flag2 = 1 if v3 is not empty and strictly increasing
flag2 = 0 otherwiseStrictly increasing means:
v3[i+1] > v3[i]for every valid i.
flag3
flag3 = 1 if v3 is not empty and strictly decreasing
flag3 = 0 otherwiseStrictly decreasing means:
v3[i+1] < v3[i]for every valid i.
Tasks
- Create project
program_2 - Define
v1 - Define
v2 - Allocate/store
v3 - Implement the matching logic
- Set
flag1 - Set
flag2 - Set
flag3 - Configure the CPU
- Run the simulation
- Inspect the pipeline
- Record performance results
Program performance
| Program | Clock cycles | Number of Instructions | CPI | IPC |
|---|---|---|---|---|
program_1 | 151 | 126 | 1.1984 | 0.8344 |
program_2 | 952 | 624 | 1.5256 | 0.6555 |
Useful formulas
CPI = Clock cycles / Number of instructionsIPC = Number of instructions / Clock cyclesTherefore:
IPC = 1 / CPINotes / observations
Exercise 3 - Benchmark performance and workload scenarios
Benchmarks
The laboratory folder includes:
calculate_pi.s
insertion_sort.sFor each program, record:
- executed instructions;
- clock cycles;
- CPI.
Also use the results already collected for:
program_1program_2
Processor frequency
1.75 kHzRaw performance measurements
| Program | Clock cycles | Instructions | CPI | Execution time (s) |
|---|---|---|---|---|
calculate_pi | 3425 | 1067 | 3.2099 | 1.957143 |
insertion_sort | 24236 | 10690 | 2.2672 | 13.849143 |
program_1 | 151 | 126 | 1.1984 | 0.086286 |
program_2 | 952 | 624 | 1.5256 | 0.544000 |
Execution-time formula
Execution time = Clock cycles / FrequencyWith:
Frequency = 1.75 kHz = 1750 cycles/sso:
Execution time = Clock cycles / 1750Initial scenario
All programs have equal execution weight:
| Program | Weight |
|---|---|
calculate_pi | 25% |
insertion_sort | 25% |
program_1 | 25% |
program_2 | 25% |
Scenario 1
| Program | Weight |
|---|---|
program_1 | 1% |
program_2 | 63% |
calculate_pi | 25% |
insertion_sort | 11% |
Scenario 2
| Program | Weight |
|---|---|
program_1 | 20% |
program_2 | 5% |
calculate_pi | 35% |
insertion_sort | 40% |
Scenario 3
| Program | Weight |
|---|---|
program_1 | 20% |
program_2 | 31.9% |
calculate_pi | 31.4% |
insertion_sort | 16.7% |
Weighted execution time
For each program:
Weighted execution time = Program execution time × Program weightThen:
Total workload time = Sum of all weighted execution timesResults
Weighted execution times in seconds. Calculations use unrounded cycle counts and weights; displayed values are rounded to six decimal places.
| Program | Initial scenario | Scenario 1 | Scenario 2 | Scenario 3 |
|---|---|---|---|---|
calculate_pi | 0.489286 | 0.489286 | 0.685000 | 0.614543 |
insertion_sort | 3.462286 | 1.523406 | 5.539657 | 2.312807 |
program_1 | 0.021571 | 0.000863 | 0.017257 | 0.017257 |
program_2 | 0.136000 | 0.342720 | 0.027200 | 0.173536 |
| TOTAL Time (@ 1.75 kHz) | 4.109143 | 2.356274 | 6.269114 | 3.118143 |
How the measurements and weights were obtained
Each of the four programs was run separately in ASE Studio. ASE Studio assembled the RISC-V code, simulated it with gem5 and its saved CPU settings, then reported executed instructions and clock cycles. CPI is cycles divided by instructions; IPC is its reciprocal. The workload scenarios were calculated afterward, not simulated as a combined workload.
A weight is a program’s share of runs in the workload. For example, Scenario 1 assigns program_2 a weight of 63%, so its weighted contribution is 0.544000 s × 0.63 = 0.342720 s. Adding all four weighted contributions gives 2.356274 s. This is the weighted average time per run; 100 runs in those proportions would take approximately 235.6274 s. Every scenario’s weights sum to 100%.
ASE Studio suppressed the long pipeline visualizations at its 3000-cycle display limit, but reported the full totals: calculate_pi 3425 cycles and insertion_sort 24236 cycles. This display limit did not truncate either completed simulation.
Pipeline observations
Use this section while inspecting ASE Studio.
program_1
Interesting instructions
Stalls / bubbles
Branch behavior
Other observations
program_2
Interesting instructions
Stalls / bubbles
Branch behavior
Other observations
Problems encountered
- The supplied calculate_pi used syscall 10. gem5’s Linux syscall handler interpreted this as unsupported fgetxattr and failed. Only the ASE copy’s ending was replaced with the lab’s End block: a0=0, a7=93, ecall. The Desktop original was left unchanged. Benchmark instruction and cycle totals refer to this compatible version.
- calculate_pi initializes its sum with fmv.s f10, f0 while f0=1.0. It therefore computes approximately pi + 1. This arithmetic issue was deliberately retained when measuring the supplied workload; no mathematical correction is included in the recorded code or timings.
- insertion_sort declares 200 words but len=50, so the measured workload sorts only the first 50 elements, in descending signed order. This supplied behavior was preserved.
What I learned
Questions
Related
Verified program outputs
- program_1 matches the provided Fibonacci template: x1=6765, x2=x5=10946, x4=21. Its comments mention printing, but the template has no print operation; no printing requirement was inferred.
- program_2 produces v3=[2, 9, -13, 16], v3_len=4 and flag1=flag2=flag3=0. Signed lb loads and blt comparisons handle negative values. Each v1 element is retained if it appears at least once in v2, preserving v1 order and repeated v1 elements; duplicate occurrences in v2 do not duplicate a match. Empty v3 sets flags to 1,0,0; a singleton sets 0,1,1.
Remaining submission work
- Complete detailed pipeline observations where required by the lab.
- Export the completed laboratory document as lab_01.pdf.
- Prepare the required submission ZIP. Pushing these notes to GitHub does not submit the laboratory assignment.
Measured program sources
See Lab 01 - Program Sources for all four assembly listings and links to their .s files. These are snapshots of the ASE sources used for the recorded runs.