Lecture 03 - Computing Virtualization Technologies and Tools
Course: Cloud Computing Technologies
Professor: Fulvio Risso, Politecnico di Torino
Source: Computing Virtualization Technologies and Tools, based on a previous version by Alex Palesandro
Length: 116 PDF pages; printed slide numbers run from 1 to 123 with gaps.
Lecture date: not stated. Number 03 follows the supplied filename.
Source numbering
References below use PDF pages, followed by printed slide numbers where helpful. They are not interchangeable in this deck. Commands and code in the slides are study examples; they have not been executed or tested here.
Lecture map
| Section | PDF pages | Printed slides | Main concepts |
|---|---|---|---|
| I. Introduction | 2-6 | 3-7 | Virtual hardware and resources |
| II. CPU virtualization | 7-36 | 8-39, with gaps | Rings, T&E, DBT, PV, hardware assistance |
| III. Memory virtualization | 37-50 | 40-53 | Shadow tables, EPT/NPT, TLB |
| IV. I/O virtualization | 51-63 | 54-66 | Emulation, PV devices, assignment, IOMMU, SR-IOV |
| V. Hypervisor architectures | 64-69 | 67-72 | Type 1, Type 2, hybrid |
| VI. Linux tools and VM mobility | 70-107 | 73-113, with gaps | QEMU/KVM, libvirt, Virtio, vhost-net, migration |
| VII. Nested virtualization | 108-114 | 115-121 | L0/L1/L2, VMCS, nested EPT |
| Takeaways and references | 115-116 | 122-123 | Isolation versus overhead; source bibliography |
What the professor covered
Virtual hardware and the virtualization contract (PDF pages 3-9)
A virtual hardware profile specifies the machine visible to the guest. It need not match the physical equipment: one physical NIC from vendor X can back two virtual NICs modeled on vendor Y. CPU, memory, and I/O all require virtualization.
The CPU discussion assumes a compatible host/guest ISA. Running a different ISA requires instruction translation or emulation. PDF page 9 introduces the Popek-Goldberg requirements: an equivalent execution environment, efficient execution of a dominant portion of instructions on the real processor, and complete VMM control of resources.
Related: Virtual Machines and Hypervisors, CPU Virtualization and Trap and Emulate.
Privilege rings and the limits of classical T&E (PDF pages 10-21)
Traditional x86 uses ring 0 for the kernel and ring 3 for applications. Early virtualization de-privileges the guest kernel. The professor contrasts a 0/1/3 arrangement with 0/3/3 ring compression, illustrating difficulties protecting both the VMM and the guest kernel.
Trap and emulate: execute ordinary guest work, trap operations requiring privileged control, emulate their intended effect, then resume the guest. Privileged and sensitive describe different properties. Classical T&E needs sensitive operations to be interceptable; traditional x86 has sensitive operations that do not reliably trap when de-privileged.
The hand drawing on PDF page 12 distinguishes app-to-OS system calls from OS-to-hardware operations. Pages 15-16 describe extra syscall handling in an older de-privileged model; the reported possible 10× penalty belongs to that scenario. Hardware assistance later lets ordinary guest syscalls remain inside the guest.
Related: CPU Virtualization and Trap and Emulate.
DBT and paravirtualization (PDF pages 22-29)
Dynamic Binary Translation (DBT) transforms guest binary basic blocks at runtime and caches translated code without changing guest source. In the example on PDF page 24 / printed slide 26, some instructions remain unchanged while a call is rewritten with explicit return-address handling and a jump through translated-code dispatch. The original VMware design combined direct execution / T&E with DBT rather than interpreting every instruction.
Paravirtualization (PV) changes the guest kernel to cooperate using hypercalls and asynchronous notifications. The application ABI remains unchanged in the model, so applications can remain unmodified. The slide’s early PV design executes the guest kernel in ring 1 and can boot directly without conventional BIOS startup.
PDF page 29 / slide 31 contrasts source-level cooperation with binary-level rewriting. Hardware and guest compatibility determine the tradeoff; historical OS/vendor examples are not current support matrices.
Related: Dynamic Binary Translation, Paravirtualization and Hypercalls, x86 Virtualization Techniques.
Hardware-assisted CPU virtualization (PDF pages 30-36)
Intel VT-x and AMD SVM supply hardware support for controlled guest execution. In the Intel model, VMX root/non-root mode is separate from ring privilege: the VMM runs in root mode, while the guest kernel can run in non-root ring 0 and guest apps in non-root ring 3.
- VM entry: VMM to guest.
- VM exit: guest to VMM, according to architectural requirements and configured controls.
- VMCS (Intel) / VMCB (AMD): guest state and execution-control structures; the mechanisms are conceptually related but not binary-compatible.
Ordinary guest work can execute directly. VM-entry/exit round trips introduce overhead. The historical table on PDF page 33 / slide 36 reports 3963 cycles for Prescott and 784 for Sandy Bridge, among other generations; it comes from a 2012 paper and is not a measurement of current CPUs.
Related: Hardware-Assisted Virtualization and VMCS.
Memory translation (PDF pages 38-50)
Virtualized memory adds an address domain:
Guest virtual address -> Guest physical address -> Machine physical addressShadow page tables compose the two mappings into a guest-virtual-to-machine-physical table used by the CPU. The VMM must synchronize it with guest changes. The brute-force approach traps CR3 writes and protects relevant guest page-table pages; the lazy approach reduces traps and synchronizes through selected invalidations and faults.
EPT/NPT supplies a separate hardware-managed guest-physical-to-machine-physical translation. This avoids much shadow-table synchronization work but makes uncached page walks more expensive. PDF page 48 / printed slide 51 gives 20 additional EPT table accesses for the illustrated four-level walk; this is not the total number of accesses on every load.
Tagged TLBs associate cached translations with a virtual-processor identifier, avoiding unconditional flushes at every entry/exit. Huge pages can improve TLB reach. Performance-gain figures cited from 2009 are workload-specific historical evidence.
Related: Memory Virtualization - Shadow Page Tables and EPT.
Device virtualization (PDF pages 52-63)
| Technique | Guest interface | Principal tradeoff |
|---|---|---|
| Device emulation | Driver for a modeled physical device | Compatibility and sharing versus emulation work |
| PV device | Dedicated cooperating guest driver | More efficient interface; needs appropriate driver |
| Direct assignment / passthrough | Driver for the assigned physical device | Low mediation overhead; dedicated resources and migration constraints |
PV devices use a guest frontend and host backend. Installing PV drivers is a narrower change than porting an entire kernel for classical PV CPU virtualization. The balloon driver is a PV-only example for reclaiming/managing guest memory.
Direct device access introduces DMA-address and isolation problems. An IOMMU remaps and constrains device DMA. SR-IOV lets supporting hardware expose virtual functions assignable to different guests, with multiplexing handled by the device.
Related: I-O Virtualization - Emulation PV and Passthrough, Memory Ballooning.
Hypervisor architectures (PDF pages 65-69)
- Type 1: runs directly on hardware; slide examples include ESXi, Xen, and Hyper-V.
- Type 2: runs on a host OS as an application; examples include VirtualBox and VMware Workstation.
- Hybrid in this lecture: integrates virtualization in the host kernel; the example is KVM.
Keep this architecture comparison separate from CPU techniques. The Windows/Hyper-V and hypervisor-coexistence remarks are slide-era examples, not verified advice for a current installation.
Related: Hypervisor Architectures.
QEMU, KVM, and management (PDF pages 71-90)
QEMU models a whole system and devices, and can translate across ISAs. User-mode emulation runs an individual foreign-ISA program; system-mode emulation provides a whole machine. KVM supplies kernel virtualization infrastructure and hardware-assisted execution on compatible ISAs.
QEMU allocates guest RAM, creates VM/vCPU handles through /dev/kvm, and runs each guest vCPU through a host thread. Linux scheduling and memory-management facilities are reused. The host sees QEMU’s threads, not each guest application thread as a separate host thread.
In the illustrated loop, QEMU calls KVM_RUN; KVM enters the guest; an exit returns to KVM, which either handles it or returns a reason to QEMU for userspace emulation. Not every hardware exit requires QEMU intervention.
PDF pages 76-77 compare same-ISA emulation with KVM acceleration and demonstrate an ARM guest on x86. The Ubuntu 20.04/21.04 image links and commands are historical examples; no images were downloaded or VMs started.
libvirt supplies management APIs and domain definitions; virsh is its CLI and virt-manager a GUI. XML records the hardware profile. Networking briefly covers bridges, NAT, physical devices, and dnsmasq; a separate lecture develops networking. Vagrant receives only a one-slide infrastructure-as-code mention.
Related: QEMU and KVM, Libvirt virsh and virt-manager.
Virtio and vhost-net (PDF pages 91-100)
Virtio standardizes cooperating guest drivers and host backends. Virtqueues communicate buffers, including scatter-gather descriptors; the lecture’s network model has TX/RX queues and an optional control queue. Scatter-gather I/O avoids requiring all buffers to be contiguous or issuing a separate operation for each buffer.
With a QEMU backend, packet processing passes through userspace before the host networking stack. vhost-net moves the network data path into the host kernel while QEMU retains setup, feature negotiation, and control responsibilities.
- TX: guest kicks the queue;
ioeventfdroutes notification toward the backend. - RX: the backend places received packets in the queue;
irqfdsupports guest interrupt notification.
The guest still sees a Virtio interface. The deck’s queue layouts, IDs, performance ratios, and worker-thread model describe its examples rather than every version/configuration.
Related: Virtio Virtqueues and vhost-net.
VM mobility (PDF pages 101-107)
Cold migration shuts down and restarts the VM, preserving persistent disks/configuration rather than execution state. Hot migration follows pause, move, resume, transferring RAM and relevant CPU/device/hypervisor state. Live migration copies state while execution continues, then completes a handoff; pages modified during copying must also be transferred.
The state diagram includes boot/configuration data, modified OS/data disks, application/OS memory, virtual-device state, and hypervisor data. Network delivery must follow the VM; the slides mention gratuitous ARP and differences between bridge/NAT setups. Remote/shared storage can reduce disk-transfer work.
Related: VM Migration - Cold Hot and Live, Virtualization Agility and VM Migration.
Nested virtualization (PDF pages 109-114)
L0 is the physical-host hypervisor, L1 the virtualized hypervisor, and L2 its guest. Use cases include OS-component testing, security tools, and honeypots.
L0 must support L1’s virtualization operations. The lecture introduces VMCS0-1, VMCS1-2, and a composed VMCS0-2. Shadow VMCS reduces selected nested-control accesses that would otherwise exit to L0.
Likewise, L0 composes EPT mappings so L2 memory reaches real physical memory. L1 table updates, faults, and invalidations must remain consistent with L0’s composed mapping.
Related: Nested Virtualization.
Study clarifications and source caveats
- PDF page 8 / slide 9 calls the original Rosetta transition “MIPS to Intel.” Apple’s announcement identifies PowerPC to Intel. This is a correction, not the professor’s wording. Apple’s 2005 announcement.
- PDF page 80 / slide 85 is schematic code. It puts RIP and general registers under a special-register setup; the x86 API separates
KVM_SET_REGS(including RIP) fromKVM_SET_SREGS. The note preserves the lifecycle rather than presenting this snippet as working C. Linux KVM API. - EPT reduces a category of synchronization overhead; “without any overhead” should not be read as zero-cost virtual memory. The same slides explicitly describe costly nested walks.
- The live-migration definition emphasizes continued service, while the final state transfer can require a short handoff. Do not infer universally zero downtime.
Questions to check understanding
- What makes a sensitive instruction different from a privileged instruction?
- Why can a guest kernel run in ring 0 without controlling the real host?
- Why does a guest syscall not necessarily cause a VM exit under hardware assistance?
- What is translated by shadow tables versus EPT/NPT?
- Why does a TLB miss become more expensive with nested paging?
- How do PV device drivers differ from whole-kernel paravirtualization?
- What problems do IOMMU and SR-IOV solve separately?
- Which component executes guest CPU code, and which manages devices and VM lifecycle?
- Why does vhost-net improve the packet path while QEMU still exists?
- Which state must cold, hot, and live migration preserve?
- Why must L0 compose L1’s VM controls and memory mappings?
Source
Fulvio Risso, Computing Virtualization Technologies and Tools, Politecnico di Torino, based on a previous version by Alex Palesandro. Original: 10 Courses/Cloud Computing Technologies/Resources/03 - Computing virtualization technologies and tools.pdf. PDF pages 1-116, printed slides 1-123 with gaps. The bibliography on PDF page 116 includes Popek and Goldberg (1974), Xen (2003), The Turtles Project (2010), nested-VM testing (2010), an EPT evaluation (2009), and an EPT walkthrough. Added explanations and corrections are labeled above.