System Design: LLM Inference Serving Platform at Scale

Design the inference serving platform for a company's LLM product: millions of requests/day, a mix of short chat replies and long streaming generations, on a fixed and expensive pool of GPUs, with a p99 time-to-first-token SLO.

What clarifying questions would you ask the interviewer?