Awesome Infra for AI › LLM Gateways & Proxies

kittors/CliRelay

⭐ 1027 Go added to this list on 2026-08-10 repository created 2026-02-27

CliRelay is a Go proxy server that consolidates AI CLI subscriptions, OAuth credentials and API keys into a single managed API layer. One endpoint on port 8317 fronts Gemini, OpenAI and Codex, Anthropic Claude, Qwen, iFlow, Kimi, Antigravity, xAI/Grok, Vertex, Bedrock, OpenCode Go, ClinePass, Ollama Cloud and any OpenAI-compatible upstream, so clients such as Claude Code, Gemini CLI, OpenAI Codex and Amp all speak to the same address. The routing layer offers round-robin or fill-first load balancing across several keys for one provider, group and path routing that binds channels into groups and restricts keys to allowed groups, and automatic failover to backup channels when quotas run out or errors appear. Text and image inputs, image-generation routing, function calling and SSE streaming are all supported. Every request is logged to PostgreSQL with timestamp, model, input, output, reasoning and cache token counts, latency, status and source channel, with full message bodies stored compressed under a separate retention policy. Pre-computed analytics cover daily trends, model distribution, hourly heatmaps and per-key statistics, alongside a health-score engine and live system stats streamed over WebSocket. The product is built for shared operation: tenants, users, roles and a fine-grained resource.action permission model decide which pages and actions each account reaches, and security-sensitive changes land in an audit log. API keys carry per-key token and request quotas, period quota resets, rate limits and reusable permission profiles binding scoped channel and model access; portal accounts group several keys under one end-user identity. The runtime stack is PostgreSQL 15+, Redis 7+ and the Ent ORM, with SQLite retained only as a migration import source. It is an enhanced fork of CLIProxyAPI.

https://github.com/kittors/CliRelay

llm-gatewayai-gatewayproxyroutingquotasmulti-tenantgoself-hosted

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.