ChatGPT — LLM Serving
PagedAttention and KV-cache memory, continuous batching, prefix-aware routing, disaggregated prefill and decode, fairness across tenants, and why VRAM is the real database.
PagedAttention and KV-cache memory, continuous batching, prefix-aware routing, disaggregated prefill and decode, fairness across tenants, and why VRAM is the real database.