diegosouzapw/OmniRoute
OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.
Awesome Infra for AI › LLM Gateways & Proxies
Olla is an efficient and high-performance proxy and load balancer specifically designed for LLM inference infrastructure. It enables intelligent routing of LLM requests across various self-hosted inference nodes, offering automatic failover and connection retry for enhanced reliability. The tool supports unified model discovery and cataloging across different providers, allowing seamless routing to available models on compatible endpoints. Olla integrates with a wide range of popular LLM inference backends such as Ollama, LM Studio, vLLM, and llama.cpp, and is designed to complement existing API gateways and orchestration platforms. It features advanced capabilities including priority-based routing, KV-cache aware sticky sessions for multi-turn conversations, smart model unification with OpenAI-compatible cross-provider routing, and dual proxy engines (Sherpa for simplicity, Olla for high performance). Health monitoring with circuit breakers, intelligent retries, self-healing model discovery, and detailed request tracking are also integral features. Olla is built for production environments, offering rate limiting, request size limits, configurable timeouts, and efficient resource usage (under 50MB RAM). It also includes Anthropic Messages API passthrough or translation and backend authentication for secure inference.
https://github.com/thushan/olla
OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.
AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...
LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.
new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...
Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...
9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.
Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.
OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.