Work / Fig. 05 · LLM inference systems · 2026

Paged KV-Cache Engine (vLLM-style)

Copy-on-write forking — ref-counted blocks shared across forked sequences.

Fig. 05

Paged KV-Cache Engine (vLLM-style)

LLM inference systems · 2026

Problem

Naïve KV-cache allocation wastes most of a GPU: memory fragments, identical prompt prefixes are recomputed, and batch scheduling stalls on the longest sequence.

Approach

Implemented vLLM's core ideas from scratch: GPU memory treated as virtual memory with fixed-size physical blocks and per-sequence page tables, ref-counted copy-on-write forking, a radix-tree prefix cache for shared prompts, and a continuous-batching scheduler moving sequences through WAITING/RUNNING/SWAPPED states.

Results

  • Copy-on-write forking

    ref-counted blocks shared across forked sequences

  • Radix-tree prefix cache

    shared prompt prefixes served without recompute

  • Continuous batching

    scheduler with swap-in/swap-out under memory pressure