Extending GVirtuS: Reducing Overhead in Remote GPU Virtualization
Authors
Nielsen, Kristian Brix Kjærsgaard ; Suneson, Sven ; Teixeira Afonso, Vasco
Term
2. semester
Education
Publication year
2026
Abstract
This thesis investigates how to reduce communication overhead in remote GPU virtualization to make GPU resources more accessible to resource-constrained IoT and edge devices. The work builds on the GVirtuS framework, which enables CUDA calls to a remote GPU through API remoting and has traditionally relied on TCP and, in HPC contexts, RDMA as communication layers. While these approaches provide high throughput under stable network conditions, they introduce significant overhead and limited robustness in wireless and unstable environments. The project therefore asks how GVirtuS’s communication overhead can be reduced through alternative GPGPU virtualization approaches, communication protocols, and related technologies. A systematic literature search, aligned with PRISMA-S principles, is used to review existing GPGPU virtualization solutions such as rCUDA, DS-CUDA, and rInfer, as well as relevant communication protocols (TCP, UDP, QUIC, RDMA). Based on this analysis, GVirtuS is extended with a QUIC-based communication layer, support for asynchronous CUDA functions, and enhanced CUDA Graph support, alongside improvements to the build system and communication components. Performance is evaluated in both emulated and real wireless test setups using matrix arithmetic, OpenPose inference, and PyTorch-based benchmarks. The results show that, under adverse network conditions, QUIC can achieve up to a 6.67× mean speed-up compared to TCP, and that combining QUIC with asynchronous communication in async-heavy applications yields up to a 6.72× mean speed-up over the previous TCP-based implementation. Furthermore, applications using CUDA Graphs with GVirtuS achieve up to a 44× mean speed-up compared to equivalent executions without CUDA Graphs, indicating that combined software improvements in the communication layer and CUDA support can make remote GPU virtualization significantly more efficient and practical for edge and IoT scenarios.
Denne afhandling undersøger, hvordan kommunikationsoverhead i fjern GPU‑virtualisering kan reduceres for at gøre GPU‑ressourcer mere tilgængelige for ressourcebegrænsede IoT- og edge-enheder. Udgangspunktet er GVirtuS‑frameworket, som muliggør CUDA‑kald til en fjern GPU via API‑remoting, og som traditionelt har benyttet TCP og i HPC‑sammenhænge RDMA som kommunikationslag. Disse løsninger fungerer godt under stabile netværksforhold, men giver væsentlig overhead og begrænset robusthed i trådløse og ustabile miljøer. Projektet formulerer derfor spørgsmålet om, hvordan GVirtuS’ kommunikationslag kan forbedres gennem alternative protokoller og relaterede teknologier. Med en systematisk litteraturgennemgang afdækkes eksisterende GPGPU‑virtualiseringsløsninger som rCUDA, DS‑CUDA og rInfer samt relevante kommunikationsprotokoller (TCP, UDP, QUIC, RDMA). På baggrund heraf udvides GVirtuS med et QUIC‑baseret kommunikationslag, understøttelse af asynkrone CUDA‑funktioner og udvidet support for CUDA Graphs, og der gennemføres forbedringer i build‑system og kommunikationskomponenter. Ydelsen evalueres i både emulerede og realistiske trådløse testmiljøer ved hjælp af matrixberegninger, OpenPose-inferens og PyTorch‑baserede benchmarks. Resultaterne viser, at QUIC under dårlige netværksbetingelser kan opnå op til 6,67× gennemsnitlig hastighedsforbedring i forhold til TCP, og at kombinationen af QUIC og asynkron kommunikation i async‑tunge applikationer kan give op til 6,72× forbedring sammenlignet med den tidligere TCP‑baserede løsning. Anvendelse af CUDA Graphs med GVirtuS giver desuden op til 44× gennemsnitlig hastighedsforbedring sammenlignet med tilsvarende eksekveringer uden CUDA Graphs, hvilket peger på, at kombinerede softwareforbedringer i kommunikationslag og CUDA‑understøttelse kan gøre fjern GPU‑virtualisering markant mere effektiv og anvendelig i edge‑ og IoT‑scenarier.
[This abstract has been generated with the help of AI directly from the project full text]
