Awesome Infra for AI › LLM Gateways & Proxies

azrtydxb/Fastllm-proxy

⭐ 108 Rust added to this list on 2026-08-24 repository created 2026-08-06

fastllm-proxy is a gateway written in Rust that fronts any number of inference backends behind a single OpenAI-compatible address, with overhead measured at well under a microsecond of work per request. Two design decisions produce that number. First, there is no I/O on the request path: role-based access control, rate limits and budgets are integer comparisons against an in-memory snapshot of the configuration, and a test in the repository fails the build if that property is lost. Second, there is no parsing on the response path: upstream frames reach the client exactly as they arrived, never deserialised and never re-encoded, so a streaming response does not cost thousands of parse cycles per second per stream. Routing is aware of how modern inference servers work. vLLM and SGLang keep a radix prefix cache, so requests sharing a system prompt are much cheaper on the node that already holds that cache; the proxy sends a shared prefix back to that node unless it is meaningfully hotter than the least-loaded one, which round-robin balancing cannot do. Rule-based and semantic routing, frontend model aliases, and eighty upstream providers plus any OpenAI-compatible endpoint are supported. Availability is treated as part of routing: a proxy that loses its control plane keeps serving from its last snapshot, per-replica health is tracked separately rather than merged, and a signal reloads the routing table without dropping a stream. Access control uses real API keys with per-principal and per-model grants, hashed credentials, and per-key budgets. LiteLLM-format configuration files are read unchanged, and a management interface is compiled into the binary.

https://github.com/azrtydxb/Fastllm-proxy

llm-gatewayllm-routeropenai-compatiblevllmsglangrbacrust

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.