Lab 01 - gem5 and ASE Studio

Course: Computer Architectures
Laboratory: 0x01
Deadline: 06/10/2026
Status: Programs and measurements completed; PDF and submission pending

Submission

Expected submission:

SURNAME_NAME_MATRICOLA_lab_01.zip

The archive must include:

  • program_1
  • program_2
  • lab_01.pdf

Important

Complete the laboratory document and export it as lab_01.pdf before creating the final submission archive.


Objective

Use ASE Studio and gem5 to:

  • write and run RISC-V assembly programs;
  • configure the simulated CPU;
  • visualize pipeline execution;
  • collect performance measurements;
  • compute CPI, IPC, and execution times;
  • compare workload scenarios.

Environment Setup

1. Open a terminal

Open a terminal on the Linux-based machine.

2. Download the simulation environment

ASE Studio is included as a Git submodule of the ase_riscv_gem5_sim repository.

Do not clone the ASE Studio repository separately.

Move to the directory where you want to install the environment and run:

git clone --branch main --recurse-submodules https://github.com/cad-polito-it/ase_riscv_gem5_sim.git
cd ase_riscv_gem5_sim

The --recurse-submodules option also downloads ASE Studio inside the parent repository.

3. Install the required tools

Labinf machine

gem5 and the RISC-V toolchain are already installed. Install only ASE Studio:

./utils/installation.sh labinf

Local virtual machine

All required tools are already installed.

Local machine

Install all required tools:

./utils/installation.sh all-ase

The complete installation may take several minutes depending on the machine and Internet connection.

4. Launch ASE Studio

Launch ASE Studio from the desktop shortcut, if available, or from the terminal:

./ase-studio.sh

5. Create a new program

In ASE Studio:

  1. Click New.
  2. Enter the project name required by the exercise.
  3. Click Create.

Example:

program_1

6. Configure the CPU

Configure the CPU according to the configuration specified in the exercise.

CPU settings used for the measurements are recorded below from the saved gem5 run commands.

7. Run the simulation

ASE Studio compiles the program and runs it with gem5 using the selected CPU configuration.

8. Analyze the pipeline

After the simulation finishes:

  1. Open the pipeline visualization.
  2. Inspect how instructions move through the pipeline stages.
  3. Use the visualization to answer the exercise questions.

9. Configure submission settings

Open Settings in ASE Studio.

In Before assignment name, enter:

SURNAME_NAME_MATRICOLA

The After assignment name field does not need to be filled unless explicitly requested.

ASE Studio will use these settings when generating the final submission filename.

10. Prepare the submission

Complete all exercises and export the finished laboratory document as:

lab_01.pdf

Then open Prepare submission in ASE Studio.

Set:

Assignment name: lab_01

Select these projects:

program_1
program_2

Under Additional files, attach:

lab_01.pdf

Finally, click Create submission.

Expected result:

SURNAME_NAME_MATRICOLA_lab_01.zip

CPU Configuration

Measurements recorded on 2026-10-03. The user confirmed that the exercise CPU settings had already been configured. The following values come from the actual saved gem5 commands, rather than an independent check of the official CPU specification.

Settingprogram_1program_2calculate_piinsertion_sort
CPU modelMinorCPUMinorCPUMinorCPUMinorCPU
Forwardingonoffonoff
Floating ALU latency (cycles)2222
Floating multiply latency (cycles)7577
Floating divide latency (cycles)8883

Common settings: integer ALU/multiply/divide latency 1 cycle; all listed arithmetic units pipelined except floating divide; direct memory; instruction/read/write latency 1 cycle; cache stalls off. The commands specify 32 KiB instruction/data caches, 64-byte cache lines, cache latency 2 and memory latency 30, while using direct-memory mode. The simulated CPU and system clocks are 1 GHz. The exercise’s 1.75 kHz is used afterward to convert measured cycles into the requested execution times; the simulator clock was not changed to 1.75 kHz.

See Lab 01 - Program Sources for the exact measured assembly and saved simulator commands for configuration evidence.


Exercise 1 - Fibonacci

Project name:

program_1

Goal

Write and run a RISC-V assembly program that calculates the Fibonacci sequence.

Starting assembly template

# Data section
 
.section .data
 
# Place here your program data. In this example
# two vector of floats, a vector of ints and a single int are defined
 
T0: .word 0x0BADC0DE
 
# Code section
 
.section .text
 
# The _start label signals the entry point of your program
# DO NOT CHANGE ITS NAME. It must be "_start", not "start",
# not "main", not "start_".
# It's "_start" with a leading underscore and
# all lowercase letters
 
.globl _start
 
_start:
 
Main:
 
    # Initialize Fibonacci variables
 
    li x1, 0        # x1 = a = first Fibonacci number (0)
 
    li x2, 1        # x2 = b = second Fibonacci number (1)
 
    li x3, 21       # x3 = count = number of terms to generate
 
    li x4, 0        # x4 = i = loop counter
 
    # Loop to generate and print remaining 20 numbers
 
    addi x4, x4, 1  # i = 1 (start from second iteration)
 
fib_loop:
 
    beq x4, x3, End # if i == count, exit loop
 
    # Calculate next Fibonacci number
 
    add x5, x1, x2  # x5 = next = a + b
 
    # Update variables for next iteration
 
    mv x1, x2       # a = b (previous second becomes first)
 
    mv x2, x5       # b = next (calculated next becomes second)
 
    # Increment counter and continue loop
 
    addi x4, x4, 1  # i++
 
    j fib_loop      # Jump back to loop start
 
End:
 
    # exit() syscall. This is needed to end the simulation gracefully
 
    li a0, 0
    li a7, 93
    ecall

Equivalent C pseudocode

/*
 * Fibonacci sequence generator in C (pseudocode)
 *
 * Generates the first 21 Fibonacci numbers using iterative approach
 */
 
int main() {
    int i;          // Loop counter
    int a = 0;      // First Fibonacci number
    int b = 1;      // Second Fibonacci number
    int next;       // Next Fibonacci number
    int count = 21; // Number of terms to generate
 
    // Generate and print remaining numbers
    for (i = 1; i < count; i++) {
        next = a + b;   // Calculate next Fibonacci number
        a = b;          // Update: previous second becomes first
        b = next;       // Update: calculated next becomes second
    }
 
    return 0;
}

Tasks

  • Create project program_1
  • Insert the Fibonacci assembly program
  • Configure the CPU
  • Run the simulation
  • Inspect the pipeline
  • Record clock cycles
  • Record executed instructions
  • Compute CPI
  • Compute IPC

Results

MetricValue
Clock cycles151
Number of instructions126
CPI1.1984
IPC0.8344

Notes / observations


Exercise 2 - Array matching and ordering flags

Project name:

program_2

Goal

Write and run a RISC-V assembly program that compares two arrays of signed 8-bit integers and builds a third array containing matching values.

Input arrays

Each array contains 10 signed 8-bit integers.

Example:

v1: .byte 2, 6, -3, 11, 9, 18, -13, 16, 5, 1
v2: .byte 4, 2, -13, 3, 9, 9, 7, 16, 4, 7

For each element of v1, check whether it appears in v2 at least once.

Store matching values in v3.

Expected result for the example:

v3: .byte 2, 9, -13, 16

Flags

Create three 8-bit unsigned flags.

flag1

flag1 = 1 if v3 is empty
flag1 = 0 otherwise

flag2

flag2 = 1 if v3 is not empty and strictly increasing
flag2 = 0 otherwise

Strictly increasing means:

v3[i+1] > v3[i]

for every valid i.

flag3

flag3 = 1 if v3 is not empty and strictly decreasing
flag3 = 0 otherwise

Strictly decreasing means:

v3[i+1] < v3[i]

for every valid i.

Tasks

  • Create project program_2
  • Define v1
  • Define v2
  • Allocate/store v3
  • Implement the matching logic
  • Set flag1
  • Set flag2
  • Set flag3
  • Configure the CPU
  • Run the simulation
  • Inspect the pipeline
  • Record performance results

Program performance

ProgramClock cyclesNumber of InstructionsCPIIPC
program_11511261.19840.8344
program_29526241.52560.6555

Useful formulas

CPI = Clock cycles / Number of instructions
IPC = Number of instructions / Clock cycles

Therefore:

IPC = 1 / CPI

Notes / observations


Exercise 3 - Benchmark performance and workload scenarios

Benchmarks

The laboratory folder includes:

calculate_pi.s
insertion_sort.s

For each program, record:

  • executed instructions;
  • clock cycles;
  • CPI.

Also use the results already collected for:

  • program_1
  • program_2

Processor frequency

1.75 kHz

Raw performance measurements

ProgramClock cyclesInstructionsCPIExecution time (s)
calculate_pi342510673.20991.957143
insertion_sort24236106902.267213.849143
program_11511261.19840.086286
program_29526241.52560.544000

Execution-time formula

Execution time = Clock cycles / Frequency

With:

Frequency = 1.75 kHz = 1750 cycles/s

so:

Execution time = Clock cycles / 1750

Initial scenario

All programs have equal execution weight:

ProgramWeight
calculate_pi25%
insertion_sort25%
program_125%
program_225%

Scenario 1

ProgramWeight
program_11%
program_263%
calculate_pi25%
insertion_sort11%

Scenario 2

ProgramWeight
program_120%
program_25%
calculate_pi35%
insertion_sort40%

Scenario 3

ProgramWeight
program_120%
program_231.9%
calculate_pi31.4%
insertion_sort16.7%

Weighted execution time

For each program:

Weighted execution time = Program execution time × Program weight

Then:

Total workload time = Sum of all weighted execution times

Results

Weighted execution times in seconds. Calculations use unrounded cycle counts and weights; displayed values are rounded to six decimal places.

ProgramInitial scenarioScenario 1Scenario 2Scenario 3
calculate_pi0.4892860.4892860.6850000.614543
insertion_sort3.4622861.5234065.5396572.312807
program_10.0215710.0008630.0172570.017257
program_20.1360000.3427200.0272000.173536
TOTAL Time (@ 1.75 kHz)4.1091432.3562746.2691143.118143

How the measurements and weights were obtained

Each of the four programs was run separately in ASE Studio. ASE Studio assembled the RISC-V code, simulated it with gem5 and its saved CPU settings, then reported executed instructions and clock cycles. CPI is cycles divided by instructions; IPC is its reciprocal. The workload scenarios were calculated afterward, not simulated as a combined workload.

A weight is a program’s share of runs in the workload. For example, Scenario 1 assigns program_2 a weight of 63%, so its weighted contribution is 0.544000 s × 0.63 = 0.342720 s. Adding all four weighted contributions gives 2.356274 s. This is the weighted average time per run; 100 runs in those proportions would take approximately 235.6274 s. Every scenario’s weights sum to 100%.

ASE Studio suppressed the long pipeline visualizations at its 3000-cycle display limit, but reported the full totals: calculate_pi 3425 cycles and insertion_sort 24236 cycles. This display limit did not truncate either completed simulation.


Pipeline observations

Use this section while inspecting ASE Studio.

program_1

Interesting instructions

Stalls / bubbles

Branch behavior

Other observations

program_2

Interesting instructions

Stalls / bubbles

Branch behavior

Other observations


Problems encountered

  • The supplied calculate_pi used syscall 10. gem5’s Linux syscall handler interpreted this as unsupported fgetxattr and failed. Only the ASE copy’s ending was replaced with the lab’s End block: a0=0, a7=93, ecall. The Desktop original was left unchanged. Benchmark instruction and cycle totals refer to this compatible version.
  • calculate_pi initializes its sum with fmv.s f10, f0 while f0=1.0. It therefore computes approximately pi + 1. This arithmetic issue was deliberately retained when measuring the supplied workload; no mathematical correction is included in the recorded code or timings.
  • insertion_sort declares 200 words but len=50, so the measured workload sorts only the first 50 elements, in descending signed order. This supplied behavior was preserved.

What I learned


Questions


Related

Verified program outputs

  • program_1 matches the provided Fibonacci template: x1=6765, x2=x5=10946, x4=21. Its comments mention printing, but the template has no print operation; no printing requirement was inferred.
  • program_2 produces v3=[2, 9, -13, 16], v3_len=4 and flag1=flag2=flag3=0. Signed lb loads and blt comparisons handle negative values. Each v1 element is retained if it appears at least once in v2, preserving v1 order and repeated v1 elements; duplicate occurrences in v2 do not duplicate a match. Empty v3 sets flags to 1,0,0; a singleton sets 0,1,1.

Remaining submission work

  • Complete detailed pipeline observations where required by the lab.
  • Export the completed laboratory document as lab_01.pdf.
  • Prepare the required submission ZIP. Pushing these notes to GitHub does not submit the laboratory assignment.

Measured program sources

See Lab 01 - Program Sources for all four assembly listings and links to their .s files. These are snapshots of the ASE sources used for the recorded runs.