Browsing: Mlops

Every major inference framework shipped prefill-decode disaggregation this year. NVIDIA built it into Dynamo. SGLang made it the default for large-scale deployments. vLLM added a KV connector API to support it natively. The consensus is forming fast: split your prefill and decode onto separate GPU pools, and throughput improves.