AI workloads are fundamentally different from the applications we’ve been running on Kubernetes for the past decade. Models consume GPUs that cost thousands of dollars per month. Agents burn through reasoning tokens around the clock. And the complexity of the stack, from model serving engines to GPU resource management, makes traditional performance tuning feel like a warm-up exercise.
Deck used in the RedHat and Akamas webinar: https://akamas.io/events/webinar-genai-optimization-kubernetes