Awesome Infra for AI › LLM Gateways & Proxies

RelayPlane/proxy

⭐ 204 TypeScript added to this list on 2026-09-14 repository created 2026-02-03

RelayPlane is a local, open-source proxy that sits between AI agents and their model providers, acting as a drop-in replacement for the Anthropic and OpenAI base URLs. It prices every request as it happens and rolls the cost up per run, agent and model in a local dashboard and a terminal cost ticker, with no SDK or code changes required beyond pointing an existing tool at the proxy. Work can be tagged at dispatch time with run and agent labels via custom headers, which Claude Code and other clients already forward, so a fan-out of many agents can be attributed to the job that spent the money. Budget caps can be set per day, hour, request, session or run, and a per-run cap returns a 429 to stop just that job instead of the whole machine; a kill switch halts all routed traffic instantly with an audit trail of what was stopped and what it saved. Opt-in anomaly detectors flag token-explosion loops, velocity spikes and repetition from a stuck agent looping the same call. Model routing sends simple work to cheap models and hard work to frontier models by a hot-reloaded complexity-tier config, without a code change, and on a 429 or provider overload RelayPlane can fail over to another provider. It supports Anthropic, OpenAI, Google Gemini, xAI, OpenRouter and local Ollama behind one base URL, runs entirely on the user's machine with credentials and traffic staying local by default, and stores its cost ledger in SQLite. It targets developers running multi-agent or fan-out agent workloads who need to see and control what each run costs.

https://github.com/RelayPlane/proxy

LLM proxycost trackingbudget capsmodel routinglocal-first

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.