Upgrade to Pro — share decks privately, control downloads, hide ads and more …

How WebAssembly Enhances Edge AI Operations

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
Avatar for Saza Saza
July 28, 2026
17

How WebAssembly Enhances Edge AI Operations

KubeCon + CloudNativeCon Japan, Japan Community Day

Avatar for Saza

Saza

July 28, 2026

Transcript

  1. Self Introduction Soichiro Ueda @saza-ku @saza_ku • A Newbie Infrastructure

    Engineer • Interest: Cloud Computing, Systems Software Ai Nozaki @ainozaki @ainno321 • Ph.D student at the University of Tokyo • Interest: Hardware Acceleration, Systems Software 2
  2. Rise of Edge AI • Edge AI runs inference directly

    on edge devices ◦ Enables low-latency local decisions ◦ Protects privacy-sensitive data Measuring advertising effectiveness at 7-Eleven [1] Biometric monitoring by SlateSafety [2] [1]https://www.sony-semicon.com/en/news/2024/2024042401.html [2]https://www.edgeimpulse.com/case-studies/slate-safety-predicting-heat-exhaustion
  3. Edge AI Requires Updates after Deployment • Edge AI follows

    a continuous lifecycle ◦ Update to newer models ◦ Update pre/postprocessing logic Support updating model via over-the-air (OTA) Collect Data Deploy Improve Model
  4. Problem of Edge AI operation • Device-specific artifacts Python package

    container NVIDIA Jetson OS: Ubuntu CPU: Arm Cortex-A
  5. Problem of Edge AI operation • Device-specific artifacts Python package

    container STM specific firmware NVIDIA Jetson STM OS: Ubuntu OS: FreeRTOS/Zephyr CPU: Arm Cortex-A CPU: Arm Cortex-M
  6. Problem of Edge AI operation • Device-specific artifacts Python package

    container STM specific firmware NVIDIA Jetson STM OS: Ubuntu OS: FreeRTOS/Zephyr CPU: Arm Cortex-A CPU: Arm Cortex-M ESP specific firmware ESP OS: FreeRTOS CPU: Xtensa
  7. The Problem Grows with Scale • Device-specific operation does not

    scale ◦ More device families, more build and deploy pipelines ◦ More model and application versions, more artifacts New model New model (C++) Device specific⚠ App Device specific⚠ firm ware
  8. Toward Cloud-like Operation for Edge AI • Containers provide a

    common deployment unit ◦ Applications share a common packaging model ◦ Orchestrators manage life cycles of applications Can we bring this operational model to edge devices?
  9. Why existing cloud techniques don’t fit? • Resource constrained devices

    cannot run containers Feasible but costs much No Linux No Linux ex. nvidia/l4t-pytorch: 5.6GB NVIDIA Jetson STM OS: Ubuntu OS: FreeRTOS/Zephyr CPU: Arm Cortex-A CPU: Arm Cortex-M ESP OS: FreeRTOS CPU: Xtensa
  10. Why existing cloud techniques don’t fit? • Container images remain

    architecture-dependent NVIDIA Jetson OS: Ubuntu CPU: Arm Cortex-A Dev env OS: Ubuntu CPU: x86-64
  11. Existing Approaches for the Edge • K3s ◦ Lightweight Kubernetes

    distribution ◦ Still requires Linux • KubeEdge ◦ Extends Kubernetes management to edge nodes ◦ Edge devices don’t typically execute workloads ▪ KubeEdge manages them as device resources
  12. Our Focus: WebAssembly • Wasm is a portable binary format

    ◦ Compiled from multiple programming languages ◦ Portable across CPUs and OSs ◦ Sandboxed for isolated execution
  13. WASI: System Interfaces for Wasm • WASI provides standardized interfaces

    for Wasm app ◦ ex. Files and streams ◦ ex. Networking Wasm App. Wasm App. WASI Wasm runtime Host kernel
  14. WASI-NN • Standardized WASI for Neural Networks ◦ Wasm application

    invokes inference through WASI-NN ◦ Inference engines remain inside the runtime Wasm App. Wasm App. WASI-NN Wasm runtime Inference Engine • • • • load() set_input() compute() get_output()
  15. WebAssembly for Edge Devices • Wasm provides a portable and

    smaller deployment unit ◦ Does not require a full Linux userspace ◦ Portable across CPU architectures 2.47GB 100MB → Wasm is a promising deployment format for edge devices
  16. How to manage IoT Devices with Kubernetes Only runs on

    Linux Container Runtime Wasm Runtime Kubelet ? Linux Embedded OS Doesn’t run on small devices Regular Kubernetes Node IoT Device
  17. How to manage IoT Devices with Kubernetes 💡 Edge server

    runs Kubelet instead Proxy Kubelet Wasm Runtime Linux Embedded OS Edge Server IoT Device
  18. Wasm Runtime: waiot • Wasm Runtime targeting small IoT devices

    (around 500KiB RAM) • Use WebAssembly Micro Runtime (WAMR) ◦ Manage the runtime ◦ Works like containerd / CRI-O ▪ containerd ↔ waiot ▪ runc ↔ WAMR bytecodealliance/wasm-micro-runtime WAMR waiot Embedded OS • Supported Platforms: ◦ ESP32 ◦ STM32 (WIP) github.com/mewz-project/waiot
  19. Proxy Kubelet: waiot-kubelet • Works as Kubelets for IoT devices

    ◦ Runs on Linux Edge Servers ◦ Cooperates with waiot waiot-kubelet • Watch Pod Resources → Make waiot run Wasm ◦ Call remote waiot, instead of local container runtime mewz-project/waiot-kubelet waiot waiot waiot
  20. How Proxy Kubelet Works control plane created → scheduled to

    device-1 waiot-kubelet waiot waiot device-1 device-2
  21. Virtual Kubelet • Library to create your own Kubelet implementation

    • Defines Interfaces that must be implemented for Kubelet ◦ e.g. CreatePod, UpdatePod, DeletePod, etc virtual-kubelet/virtual-kubelet
  22. Future Work: NPU/GPU Support Wasm App. Wasm App. WASI-NN Wasm

    runtime ARM TrustZone (TEE) Secure World Inference Engine
  23. Conclusion • How to manage Edge AI models? ◦ Frequent

    Updates ◦ Constrained Resources → Wasm meets the Edge AI requirements • How to manage Wasm on IoT devices? → Make devices Kubernetes nodes ◦ waiot: Wasm Runtime ◦ waiot-kubelet: Proxy Kubelet • Future Works: AI Inference with GPU/NPU, TrustZone