Upgrade to Pro — share decks privately, control downloads, hide ads and more …

The GenAI Optimization Triangle: Balancing Cost...

Sponsored · SiteGround - Reliable hosting with speed, security, and support you can count on.

The GenAI Optimization Triangle: Balancing Cost, Latency, and Accuracy on Kubernetes

AI workloads are fundamentally different from the applications we’ve been running on Kubernetes for the past decade. Models consume GPUs that cost thousands of dollars per month. Agents burn through reasoning tokens around the clock. And the complexity of the stack, from model serving engines to GPU resource management, makes traditional performance tuning feel like a warm-up exercise.

Deck used in the RedHat and Akamas webinar: https://akamas.io/events/webinar-genai-optimization-kubernetes

Avatar for Stefano Doni

Stefano Doni

September 10, 2026

More Decks by Stefano Doni

Other Decks in Technology

Transcript

  1. The GenAI Optimization Triangle Balancing Cost, Latency, and Accuracy on

    Kubernetes © 2026 Akamas • All Rights Reserved • Confidential
  2. Speakers Stefano doni Daniele Zonca Roland Huß CTO & Co-Founder

    Distinguished Engineer Distinguished Engineer Akamas & Chief Architect & Architect Red Hat Red Hat © 2026 Akamas • All Rights Reserved • Confidential
  3. Agenda • Intro • Agentic AI Systems on Kubernetes •

    The GenAI Optimization Triangle • Inference optimization: from GPU to the full stack • AI for AI: the next optimization frontier • Q&A © 2026 Akamas • All Rights Reserved • Confidential
  4. Anatomy of an Agentic AI System © 2026 Akamas •

    All Rights Reserved • Confidential
  5. Anatomy of an Agentic AI System © 2026 Akamas •

    All Rights Reserved • Confidential
  6. The GenAI Optimization Triangle  Cost © 2026 Akamas •

    All Rights Reserved • Confidential   Latency Accuracy
  7. CPU vs GPU architecture • Generalist • 10s of threads

    • Optimized for latency © 2026 Akamas • All Rights Reserved • Confidential Source: NVIDIA • Specialist • 1000s of threads • Optimized for throughput
  8. How does LLM inference work? Source: RedHat TPOT GPU compute

    bound © 2026 Akamas • All Rights Reserved • Confidential TPOT TPOT GPU memory bandwidth bound
  9. AI apps can have diverse SLOs Source: Hao AI Lab

    Real time vs batch: latency vs throughput © 2026 Akamas • All Rights Reserved • Confidential
  10. What drives your AI infra cost and performance? (hint: your

    workload) © 2026 Akamas • All Rights Reserved • Confidential Source: NVIDIA
  11. AI Factory: efficiency vs performance tradeoff different configs © 2026

    Akamas • All Rights Reserved • Confidential
  12. What if you can extract 2x TPS from your GPUs?

    optimal config 2X default config © 2026 Akamas • All Rights Reserved • Confidential
  13. AI inference is a full-stack optimization problem Application & Workload

    • Workload type & SLOs • RAG configs (e.g. VectorDB Intelligent Routing • Automatic model selection • KV-cache aware routing Model Inference Server Kubernetes © 2026 Akamas • All Rights Reserved • Confidential • Model size & type • Quantization • Parallelism/distribution, KV cache • Aggregation vs disaggregation • Batching and concurrency • Pod requests & limits • HPA scaling policy • Replicas GPU/Accelerator • CUDA kernels configs • Library settings (e.g. NICCL • GPU sizing & sharing HW Instance • Cloud instance/GPU type • Number of instances • Storage & network params
  14. Multi GPUs - Data Parallelism © 2026 Akamas • All

    Rights Reserved • Confidential
  15. AI factory ROI == GPU utilization Source: VentureBeat 5% CPU

    used, out of 60.000 provisioned $15M/y) Allocatable CPUs Requested CPUs Used CPUs 1 AM © 2026 Akamas • All Rights Reserved • Confidential 5 AM 4 PM 11 PM Source: Akamas real customer European enterprise company), aggregate data for 1 day of the entire K8s footprint
  16. Optimizing AI systems — with AI Autonomous Tuning AI SRE

    Security & vulnerability AI/ML finds optimal full-stack configurations no human can tune manually at scale. Maximize throughput, minimize latency and GPU costs. AI helps you fix incidents faster. LLM-powered assistants correlate signals across metrics, logs, and traces. Cut MTTR and let SREs focus on resolution. AI agents autonomously scan codebases, identify vulnerabilities, and propose fixes — at a scale and depth no human security team can cover. © 2026 Akamas • All Rights Reserved • Confidential
  17. Get in touch with our speakers Stefano doni Daniele Zonca

    Roland Huß Akamas Red Hat Red Hat linkedin.com/in/stefanodoni/ linkedin.com/in/daniele-zon linkedin.com/in/ro14nd/ ca-9867807 © 2026 Akamas • All Rights Reserved • Confidential
  18. Contacts LinkedIn X @akamas AkamasLabs Email [email protected] Website © 2026

    Akamas • All Rights Reserved • Confidential akamas.io