Awesome Infra for AI › LLM Gateways & Proxies

atopos31/llmio

⭐ 317 TypeScript repository created 2025-08-06

LLMIO is a robust, Go-based LLM load-balancing gateway designed to streamline the integration and management of multiple large language models (LLMs) from different providers. It offers a unified REST API compatible with OpenAI Chat Completions, Anthropic Messages, and Gemini Native, supporting both streaming and non-streaming passthrough. Key features include weighted scheduling with two strategies (random by weight, priority by weight) for intelligent request routing based on tool calling, structured output, and multimodal capabilities. The platform boasts an intuitive Admin Web UI built with React, TypeScript, Tailwind, and Vite, enabling easy configuration of providers, models, associations, and real-time monitoring of logs and metrics. LLMIO also incorporates essential operational features such as built-in rate limiting with fallback mechanisms, provider connectivity checks for fault isolation, and local persistence for configurations and request logs using SQLite. A standout feature is its comprehensive observability, tracking every request with TraceID, detailed latency breakdowns (proxy, first-chunk, completion time), TPS, and token usage (input, cached, output). It calculates per-request costs based on configurable per-million-token prices, offering deep insights into LLM usage and expenditure. Additionally, it supports session tracking, allowing users to tag and filter logs by session ID for enhanced debugging and analysis.

https://github.com/atopos31/llmio

LLM gatewayload balancingobservabilitycost trackingOpenAIAnthropicGeminiAPI proxyLLM operationsGoGolang

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.