Awesome Infra for AI › LLM Gateways & Proxies

higress-group/higress

⭐ 9493 Go repository created 2022-10-27

Higress is an open-source, cloud-native AI gateway built upon the foundations of Istio and Envoy. It is specifically designed to manage and optimize traffic for AI/ML inference workloads, particularly focusing on Large Language Models (LLMs) and AI agents. Higress acts as a unified entry point, providing essential "Ops for AI" capabilities including observability, multi-model load balancing, token rate limiting, and caching for interactions with various LLM providers, both domestic and international. A key feature is its support for hosting Model Context Protocol (MCP) servers through a flexible plugin mechanism. This enables AI agents to seamlessly integrate and call diverse tools and services. Higress simplifies the conversion of OpenAPI specifications into remote MCP servers using the `openapi-to-mcp` tool, offering benefits like unified authentication/authorization, fine-grained rate limiting, comprehensive audit logs, and rich observability for tool calls. While it can function as a general Kubernetes ingress controller, its core value proposition lies in its AI-specific enhancements, enabling dynamic updates without disruption and secure, governed access to AI services. Originally developed at Alibaba, Higress supports critical AI applications within Alibaba Cloud, ensuring high availability and robust management for enterprise AI deployments.

https://github.com/higress-group/higress

AI gatewayLLM gatewayAPI gatewayLLM proxyAI agentsMCP servertraffic managementobservabilityKubernetescloud-nativeEnvoyIstio

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.