– Language models, image generators… – Require specialized inference approaches Different resource requirements – Often have large memory footprints and computational demands Importance of efficient inference – Crucial for generating high-quality outputs in real-time
serving software – Difficulty in keeping up with frequent releases of improved model serving software Challenges in deploying model + model serving software as container images – Difficulty in adapting to fast-paced changes in the ecosystem Complexity of integrating model serving software optimized for various AI accelerator hardware with models – Challenges in managing optimized combinations of software and models for each hardware Complexity of managing the interoperability matrix between various model serving software and supported models – Difficulty in managing and testing the matrix of supported model and software combinations
serving software, providing automatic conversion and integration – Enabling independent updates and management of models and serving software – Providing unified management and abstraction for various AI accelerator hardware
of inference Handle inference at various scales – From NVIDIA Jetson Nano & RPi4 to 2000 node clusters Catches both efficient resource utilization and automatic scaling
Leveraging the advancements in AI training for inference – Can the same techniques optimized for training large models be used for inference? Goals – Applying best practices and techniques from training to optimize inference performance – Ensuring seamless integration between training and inference pipelines
24.03 (Official) – Model session management: Sokovan-controlled inference – Model traffic: Backend.AI AppProxy v4 – Model fine-tuning pipeline: FastTrack Flexibility, performance, and ease of use – Dynamic inference: combining models and inference engines based on runtime requests – vLLM, TensorRT-LLM+Triton, ONNXruntime as pre-built model engine – Automatic model loading, dynamic batching, and request prioritization
for Inference – Combination of Backend.AI Model Player, model storage, and pre-defined models – Scalability and performance benefits of PALI Model Player Hugging Face NVIDIA NIM ION Model Store
appliance with built-in PALI – Performant AI Launcher for Inference on Portable AI Landing Infrastructure – Scalability through easy connection of multiple units – Optimized for AI workloads, offering high performance and low latency Partners More to come! PALI PALI with NVIDIA GH200 Kyocera Mirai Instant.AI
platform for LANGuage models – Ready-to-use setup for inference and fine-tuning – Simplified deployment and management of language models – Helmsman ✓ NLP interface for your model finetuning – Talkativot ✓ Chat with your model, simplified
for PALANG model fine-tuning & controlling ✓ Translate user intents into actionable commands ✓ Autonomously handle complex workflows ✓ Specific Use Case: ◦ Operating container based session creation with allocated resources ◦ Executing code on user behalf at remote session ◦ Launching LLM fine-tuning ◦ Deploying fine-tuned model for inference – All done with text prompt! Helmsman - person responsible for steering and controlling a direction of the ship
What You Get – Lablup’s goal: Translating AI ideas into reality Seamless and intuitive experience for deploying and utilizing AI models Empowering users to bring their AI visions to life effortlessly