diegosouzapw/OmniRoute
OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.
Awesome Infra for AI › LLM Gateways & Proxies
fastllm-proxy is a gateway written in Rust that fronts any number of inference backends behind a single OpenAI-compatible address, with overhead measured at well under a microsecond of work per request. Two design decisions produce that number. First, there is no I/O on the request path: role-based access control, rate limits and budgets are integer comparisons against an in-memory snapshot of the configuration, and a test in the repository fails the build if that property is lost. Second, there is no parsing on the response path: upstream frames reach the client exactly as they arrived, never deserialised and never re-encoded, so a streaming response does not cost thousands of parse cycles per second per stream. Routing is aware of how modern inference servers work. vLLM and SGLang keep a radix prefix cache, so requests sharing a system prompt are much cheaper on the node that already holds that cache; the proxy sends a shared prefix back to that node unless it is meaningfully hotter than the least-loaded one, which round-robin balancing cannot do. Rule-based and semantic routing, frontend model aliases, and eighty upstream providers plus any OpenAI-compatible endpoint are supported. Availability is treated as part of routing: a proxy that loses its control plane keeps serving from its last snapshot, per-replica health is tracked separately rather than merged, and a signal reloads the routing table without dropping a stream. Access control uses real API keys with per-principal and per-model grants, hashed credentials, and per-key budgets. LiteLLM-format configuration files are read unchanged, and a management interface is compiled into the binary.
https://github.com/azrtydxb/Fastllm-proxy
OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.
AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...
LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.
new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...
Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...
9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.
Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.
OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.