diegosouzapw/OmniRoute
OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.
Awesome Infra for AI › LLM Gateways & Proxies
OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.
AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...
LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.
new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...
Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...
9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.
Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.
OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.
Casdoor is an open-source, "AI-first" Identity and Access Management (IAM) and Model Context Protocol (MCP) gateway, providing authentication and authorization for AI applications and agents.
Portkey AI Gateway is a fast, open-source AI gateway designed for routing requests to over 1,600 LLMs, featuring integrated guardrails, automatic retries, and load balancing for reliable and secure...
Higress is a cloud-native AI gateway based on Istio and Envoy, providing unified management, observability, and traffic control for LLM APIs and Model Context Protocol (MCP) servers.
CoAI.Dev is a next-generation, multi-tenant LLM gateway and AIGC solution offering unified API access, load balancing, cost management, and various AI application features for over 200 models from ...
Bifrost is a high-performance AI gateway that unifies access to over 23 providers through a single OpenAI-compatible API, offering features like automatic failover, load balancing, semantic caching...
Open-source LLM gateway that routes agent and app requests to 300+ models through one OpenAI-compatible endpoint, with cost tracking, fallback and self-healing.
Plano is an AI-native proxy and data plane for agentic applications, providing built-in orchestration, safety, observability, and intelligent LLM routing to simplify the production deployment of AI...
Self-hosted Go gateway that unifies multiple AI providers and subscription accounts behind one endpoint, with credential scheduling, failover, logging and cost tracking.
vLLM Semantic Router is an intelligent routing system designed for managing and orchestrating diverse AI/ML models (mixture-of-models) across various environments, focusing on efficiency, safety, a...
Self-hosted or hosted Go proxy that routes each LLM request across Anthropic, OpenAI, Gemini and OpenAI-compatible providers using an on-box embedding-based scorer to cut model costs.
Agentgateway is an open-source proxy offering unified connectivity and governance for AI agents and LLM providers, encompassing security, observability, and advanced traffic management features.
Octelium is a self-hosted zero-trust secure access platform operating as a ZTNA, VPN, API/AI/MCP gateway, and PaaS, with specific features for AI/LLM gateway functionality.
CPA Manager Plus is a self-hosted dashboard for monitoring AI gateway traffic, providing real-time analysis of requests, costs, failures, quotas, and account health for OpenAI-compatible and CPA/CL...
GoClaw is a multi-tenant AI agent platform and gateway built in Go, enabling the deployment and orchestration of AI agent teams with extensive LLM provider support, sophisticated memory management,...
WindsurfAPI is an OpenAI and Anthropic compatible API proxy that translates requests to Windsurf's internal gRPC protocol, providing access to over 100 LLM models with account pooling, rate limitin...
Self-hosted unified API in front of agent harnesses such as Codex, Claude Code and Hermes, implementing the Unified Harness Protocol with sessions, streaming, files and cancellation.
CCS is a multi-provider profile and runtime manager for various AI models and APIs, enabling seamless switching between Claude, Gemini, Copilot, OpenRouter, and local models without configuration o...
Octopus is a self-hosted LLM API aggregation and load balancing service that provides a unified gateway for multiple LLM providers, intelligent routing, and analytics for cost and usage tracking.
Bionic is an on-premise, secure, and scalable LLM gateway and RAG platform designed to replace ChatGPT while maintaining data confidentiality and offering advanced features like AI assistants, toke...
OpenRelay is an AI model router and proxy that unifies numerous free and paid AI model quotas into a single local endpoint, enabling their use across various AI tools and IDEs.
Envoy-powered open source control plane for AI and agent traffic, giving one OpenAI-compatible API across model providers and self-hosted inference and MCP servers.
Self-hosted OpenAI-compatible LLM router that fronts more than a hundred models with bring-your-own-key access, streaming, automatic model selection and an optional hosted fallback.
Paddler is an open-source LLM/VLM load balancer and serving platform for self-hosting and scaling models, built around llama.cpp for efficient inference with dynamic model swapping and observability.
LLM Gateway is an open-source API gateway for Large Language Models, providing unified access, API key management, usage analytics, multi-provider routing, and performance monitoring for various LL...
Paritok is an open-source, non-destructive compression model and gateway purpose-built for coding agents, significantly reducing LLM input token costs while maintaining performance on SWE-bench.
TokenHub is an enterprise AI gateway offering role-based access, API routing, usage analytics, and a model catalog for managing various LLM providers.
GoModel is a fast, lightweight AI gateway written in Go, providing a unified OpenAI-compatible API for various LLM providers with observability, guardrails, streaming, and cost tracking.
A small AI gateway acting as an OpenAI and Anthropic-compatible proxy for GitHub Copilot, Codex, and other third-party AI providers, enabling unified access and management.
ThinkWatch Lite is a desktop gateway app that connects Claude Code, Codex and other AI CLI clients to multiple upstream model providers through one configuration, adding routing, failover, cost tracking and outbound credential redaction.
Self-hosted multi-provider AI gateway that fronts Gemini, Claude, Codex, Qwen and OpenAI-compatible upstreams behind one endpoint, adding routing groups, failover, request logging, quotas and multi-tenant governance.
EnterpriseAgentFramework is a Java/Spring Boot platform for registering, governing, orchestrating, and exposing enterprise APIs as AI capabilities for agents, focusing on production-grade AI operat...
ThinkWatch is an enterprise-grade AI gateway for secure, audited, and governed access to AI APIs and Multi-Cloud Provider (MCP) tools, providing unified proxying, RBAC, rate limiting, and cost trac...
NadirClaw is an open-source LLM router and AI cost optimizer that intelligently routes prompts to different language models based on complexity, reducing API costs by 40-70% through an OpenAI-compa...
Local gateway and dashboard that routes every AI coding agent on a machine through one policy layer, tracking spend, permissions and provider authenticity across 44 model vendors.
Lynkr is an HTTP proxy CLI tool designed to optimize interactions with LLMs, particularly for AI coding assistants, by compressing tokens, managing caching, and routing requests for efficiency and ...
Adaline Gateway is a fully local, production-grade SDK providing a unified interface for calling over 300+ LLMs with built-in features like batching, retries, caching, callbacks, and OpenTelemetry ...
Local OpenAI- and Anthropic-compatible proxy that exposes a Claude subscription to other AI tools, with session-affinity routing, multi-account pooling and drift detection for long agent runs.
Otari is an open-source, OpenAI-compatible LLM gateway that centralizes routing, authentication, budget enforcement, and usage tracking for over 40 AI providers.
gemini-web2api-go is a self-hosted Go proxy that turns the free Gemini web app into an OpenAI-compatible /v1/chat/completions API, with cookie and proxy pools, rate limiting and an admin dashboard.
GPTRouter is an AI model gateway for managing multiple LLMs and image models, providing universal API access, smart fallbacks, automatic retries, and reduced latency for reliable AI application per...
Sagify simplifies LLM and ML model deployment, management, and inference on AWS SageMaker, featuring an LLM Gateway for unified access to various large language models.
Nexus is an AI gateway that unifies access to multiple LLM providers and Model Context Protocol (MCP) servers, offering robust routing, security, and governance for AI stacks.
Kiji Privacy Proxy is an intelligent privacy layer for AI APIs that automatically detects and masks personally identifiable information (PII) in requests to AI services.
ccLoad is an AI API gateway that provides smart routing, automatic failover, exponential cooldown, multi-URL scheduling, real-time monitoring, and cost control for various LLM APIs.
Misceo is a local Anthropic-compatible gateway that sends eligible agent traffic to a cheaper model, checks the produced answer, and escalates to a stronger backend when a structural gate or an LLM judge rejects it.
ORBIT is a self-hosted AI gateway and retrieval-adapter layer designed for private, multi-model RAG applications, offering secure inference, data retrieval, and agentic tool-calling capabilities.
ccproxy is a CLI-based transparent network interceptor and proxy for LLM clients, enabling cross-provider routing, request/response transformation, and custom hooks for various large language models.
Privacy Filter is a Go-based LLM privacy gateway that redacts PII and secrets from text with millisecond latency, specifically designed to ensure data privacy before reaching large language models.
LLMIO is a Go-based LLM load-balancing gateway providing a unified API, weighted scheduling, observability, and an admin UI for managing various LLM providers.
Olla is a high-performance, lightweight proxy and load balancer for LLM infrastructure, providing intelligent routing, automatic failover, and unified model discovery across diverse inference backe...
digiRunner is an enterprise-grade API Gateway that acts as a unified control plane for both microservices and AI services, providing governance, cost control, and prompt management for LLMs.
Ferro Labs AI Gateway is a high-performance Go-native LLM gateway for routing requests across 30+ providers with features like caching, guardrails, A/B testing, and cost controls.
1flowbase is an open-source virtual model gateway that allows users to build multi-model workflows, publish them as OpenAI/Claude-compatible endpoints, and gain visibility into trace, token, latenc...
GPROXY is a Rust-based, high-performance, multi-provider LLM proxy server that unifies OpenAI, Claude, and Gemini-style APIs, offering multi-tenant authorization, rate limiting, quota management, a...
BitRouter is an open-source, local-first LLM router built in Rust that optimizes AI agent performance and cost by dynamically routing requests to the most appropriate LLM, supporting multiple provi...
Traceloop Hub is a high-performance, OpenTelemetry-based LLM gateway written in Rust, centralizing control and tracing of LLM calls across multiple providers with built-in observability.
A self-hosted LLM gateway in one Go binary that speaks four wire protocols, translates between them, fails over across providers, rotates upstream keys and ships a multi-user admin console.
Local proxy for AI agent traffic that prices every request in real time, rolls costs up per run and agent, and lets teams cap or kill runaway spend before it drains a budget.
Qwen Gate is a self-hosted, OpenAI-compatible API gateway that allows users to access Qwen AI models (via chat.qwen.ai) for free, featuring multi-account rotation, tool calling, and a web dashboard.
Nyro is a self-hosted AI gateway that translates protocols between different AI tools and model providers, enabling interoperability and flexible model routing without code changes.
A Rust-native AI gateway from the creators of Apache APISIX: one OpenAI-compatible API in front of every provider, with routing, failover, guardrails, caching and observability in a single binary.
Busbar is a self-hosted LLM gateway implemented in Rust that provides multi-vendor failover, load balancing, protocol translation, and governance for AI applications.
LM-Proxy is a lightweight, OpenAI-compatible HTTP LLM proxy/gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch), supporting real-time streaming, API key management, and dynami...
SmarterRouter is an intelligent LLM gateway and VRAM-aware router that profiles models, aggregates benchmarks, and automatically routes queries to the best available LLM, supporting local and exter...
Clipal is a local LLM API gateway and reverse proxy designed for developer productivity, offering unified access, failover, and key management for AI coding assistants like Claude Code, Codex CLI, ...
TiyGate is an open-source AI gateway built in Rust for managing and routing LLM services across multiple providers with features for high availability, logging, and usage analytics.
KeiRouter is a self-hostable, blazing-fast AI gateway that acts as a smart middleman for LLM API calls, providing intelligent routing, caching, cost control, and security guardrails.
Open Bias is an open-source reliability harness that acts as a proxy between applications and LLM providers to enforce runtime rules and policies, preventing off-policy behavior.
VoidLLM is a privacy-first, self-hosted LLM proxy and AI gateway designed for teams, offering features like load balancing, multi-provider routing, API key management, usage tracking, and rate limi...
Edgee is an open-source, Rust-powered LLM gateway that optimizes AI traffic through token compression, multi-provider routing, and real-time observability.
A FastAPI proxy that translates OpenAI- and Anthropic-compatible API requests to the GigaChat API, enabling seamless integration of GigaChat with existing LLM applications.
Self-hosted, OpenAI-compatible LLM gateway that gives each project one /v1 endpoint and API key while the operator manages upstream accounts, routing chains and fallbacks in one place.
reShapr is an open-source, no-code MCP Server that transforms traditional REST, GraphQL, and gRPC APIs into LLM-friendly tools, optimizing context windows and enabling AI-native API access.
ollamaMQ is a high-performance, asynchronous proxy and load balancer for Ollama and LM Studio APIs, providing multi-backend load balancing, fair-share queuing, model-aware routing, and a real-time ...
Nexus is an intelligent multi-LLM router providing task-aware model selection, cost optimization, and production safety controls via a drop-in OpenAI-compatible API.
CLI, Python library and local proxy that pools free and keyless tiers from 22 LLM providers behind one OpenAI-compatible endpoint with automatic failover.
CloudForge AI is a distributed LLM gateway that orchestrates multi-provider inference with intelligent routing, cost governance, and security for high-availability production environments.
Rust LLM router presenting one OpenAI-compatible endpoint in front of eighty providers and local vLLM or SGLang backends, with cache-affinity routing, RBAC, budgets and no I/O on the request path.
Local gateway that puts Claude Code, Cursor, Codex and other agent CLIs behind one endpoint, tracing usage and routing requests across local models, quotas and paid APIs.
AI Worker Proxy is a Cloudflare Workers-hosted LLM gateway exposing OpenAI- and Anthropic-compatible endpoints that routes requests across multiple providers (OpenAI, Claude, Gemini, local models) with API key rotation and automatic failover.