System Design: OpenAI/Gemini Platform

You run an LLM platform like OpenAI/Gemini. Provide API for completions/chat, fine-tuning, moderation, usage tracking, billing, and safety with global latency under 300 ms.

What clarifying questions would you ask the interviewer?