Paged KV-Cache Engine (vLLM-style)
LLM inference systems · 2026Problem
Naïve KV-cache allocation wastes most of a GPU: memory fragments, identical prompt prefixes are recomputed, and batch scheduling stalls on the longest sequence.
Approach
Implemented vLLM's core ideas from scratch: GPU memory treated as virtual memory with fixed-size physical blocks and per-sequence page tables, ref-counted copy-on-write forking, a radix-tree prefix cache for shared prompts, and a continuous-batching scheduler moving sequences through WAITING/RUNNING/SWAPPED states.
Results
Copy-on-write forking
ref-counted blocks shared across forked sequences
Radix-tree prefix cache
shared prompt prefixes served without recompute
Continuous batching
scheduler with swap-in/swap-out under memory pressure