Awesome Infra for AI › LLM Gateways & Proxies

Chleba/ollamaMQ

⭐ 127 Rust repository created 2026-02-27

`ollamaMQ` is a robust and high-performance asynchronous message queue dispatcher and load balancer designed specifically for managing inference requests to Ollama and LM Studio API instances. Acting as a smart proxy, it efficiently distributes incoming requests from multiple users across several backend models. Its core features include multi-backend load balancing using a combination of least-connections and round-robin strategies, automatically detecting and routing requests to compatible Ollama or OpenAI-compatible backends. The proxy also implements model-aware routing, ensuring requests are sent only to instances that have the required model loaded, preventing errors and optimizing resource usage. Key capabilities include parallel processing of requests to maximize throughput, continuous backend health checks to maintain reliability, and a sophisticated per-user queuing system with fair-share scheduling to prevent resource monopolization. It supports VIP and Boost modes for priority users and provides transparent header forwarding. A real-time Text-based User Interface (TUI) dashboard offers live monitoring of backend health, active requests, queue depths, and overall throughput. Built in Rust with `tokio` and `axum`, it offers an efficient and concurrent solution for managing LLM inference workloads, supporting various Ollama and OpenAI-compatible API endpoints.

https://github.com/Chleba/ollamaMQ

ollamaproxyllm gatewayload balancingfair-sharequeuingtuirustopenai-compatiblemodel-aware routinginference serving

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.