Awesome Infra for AI › Weekly › 2026-07-13

2026-07-13

7 projects added

AI Deployment on Kubernetes

pmady/keda-gpu-scaler

A KEDA External Scaler that enables Kubernetes to autoscale GPU workloads by directly using NVIDIA NVML metrics, supporting scale-to-zero for AI inference servers like vLLM and Triton.

LLM Gateways & Proxies

mozilla-ai/otari

Otari is an open-source, OpenAI-compatible LLM gateway that centralizes routing, authentication, budget enforcement, and usage tracking for over 40 AI providers.

edgee-ai/edgee

Edgee is an open-source, Rust-powered LLM gateway that optimizes AI traffic through token compression, multi-provider routing, and real-time observability.

Model Serving Frameworks

xLLM-AI/xllm

xLLM is a high-performance LLM inference engine optimized for diverse AI accelerators, particularly Chinese hardware, focusing on efficient, low-latency, and high-throughput model serving.

muxi-ai/muxi

MUXI is an open-source AI application server providing production infrastructure for deploying and operating AI agents with built-in orchestration, memory, observability, and scaling capabilities.

zhongkaifu/TensorSharp

TensorSharp is a native .NET inference engine specifically designed for serving GGUF large language models (LLMs) and DiffusionGemma-style text-diffusion models, offering console, web-based, and Op...

Workflow Orchestration for AI

jigjoy-ai/mozaik

Mozaik is a TypeScript runtime for building self-organizing AI agents, enabling them to communicate, coordinate, and adapt at runtime rather than relying on fixed workflows.

Newer issue Older issue