Awesome Infra for AI › AI Deployment on Kubernetes
pmady/keda-gpu-scaler
KEDA GPU Scaler is an external gRPC scaler for Kubernetes External ScaledObject (KEDA), designed specifically for autoscaling GPU-based AI/ML workloads. It addresses the limitation of standard Kubernetes HPA, which cannot monitor GPU utilization. Unlike traditional solutions that rely on `dcgm-exporter -> Prometheus -> KEDA` causing latency, this project directly reads NVIDIA GPU metrics (e.g., GPU utilization, memory usage, temperature, power draw, PCIe/NVLink throughput) from NVML C-bindings using `go-nvml`. This direct approach reduces latency, enabling faster and more efficient scaling decisions, including aggressive scale-to-zero for cost optimization. The scaler runs as a DaemonSet on GPU-enabled nodes, providing node-level hardware access, and integrates with KEDA via a gRPC interface. It supports various inference servers like vLLM and NVIDIA Triton, offering pre-built scaling profiles for common AI use cases (e.g., vLLM inference, Triton inference, training, batch), as well as custom configuration options for fine-grained control over metric types and thresholds. Additionally, it offers optional Prometheus-compatible metrics for monitoring the scaler and the GPU fleet's health, though scaling functions independently. This tool is crucial for efficiently managing and cost-optimizing AI inference deployments on Kubernetes.