ChatGPT
PagedAttention and KV-cache memory, continuous batching, prefix-aware routing, disaggregated prefill and decode, fairness across tenants, and why VRAM is the real database.
Serving large models: the GPU is the database, VRAM is the disk, and the scheduler is the whole system.
PagedAttention and KV-cache memory, continuous batching, prefix-aware routing, disaggregated prefill and decode, fairness across tenants, and why VRAM is the real database.