System Design: LLM Inference Serving Platform at Scale
Design the inference serving platform for a company's LLM product: millions of requests/day, a mix of short chat replies and long streaming generations, on a fixed and expensive pool of GPUs, with a p99 time-to-first-token SLO.