Awesome Infra for AI › LLM Gateways & Proxies

nyroway/nyro

⭐ 194 Rust repository created 2026-02-07

Nyro is a local AI gateway designed to sit between AI tools and model providers, offering on-the-fly protocol translation. This allows tools like Claude Code, Codex CLI, Gemini CLI, and any client utilizing OpenAI, Anthropic, or Gemini SDKs to connect to any backend model without requiring code modifications. It supports ingress protocols including OpenAI (Chat Completions + Responses API), Anthropic Messages, and Gemini GenerateContent, and can route to any OpenAI-compatible, Anthropic, or Gemini upstream. Key features include streaming passthrough with cross-protocol format conversion, `` tag parsing, tool call normalization, and both chat and embedding route types. The gateway provides robust routing capabilities, including exact match routing on `virtual_model` names, multi-target routing with weighted load balancing or priority-based failover, and health-aware failover. It also incorporates caching mechanisms for both exact and semantic matches to reuse responses for identical or similar queries. Nyro automatically detects model capabilities, supports dynamic model discovery, and includes security features like independent proxy and admin bearer tokens, default-deny authorization, and per-key quotas (RPM/RPD/TPM/TPD). Management is facilitated through a UI for providers, routes, and API keys, alongside request logs, usage charts, and configuration import/export. It offers deep integration with popular AI coding tools for one-click configuration synchronization and is deployable as a desktop app (macOS, Windows, Linux) or a standalone server binary.

https://github.com/nyroway/nyro

AI gatewayLLM proxyprotocol translationmodel routingAI securityAI cachingself-hosted AILLM managementinference gateway

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.