Awesome Infra for AI › LLM Gateways & Proxies

Continuum-AI-Corp/OrcaRouter-Lite

⭐ 1722 Python added to this list on 2026-08-24 repository created 2026-05-03

OrcaRouter Lite is the open-source single-workspace edition of the OrcaRouter service, packaged as a self-hosted router that speaks the OpenAI chat completions API. Applications point any OpenAI client library at the local endpoint and keep using the same request and response shapes, while the router decides which upstream provider serves each call. Provider credentials stay on the machine that runs the container, so keys are never handed to a third party, and a catalogue of more than a hundred models across the major providers is addressable by name. Passing the model auto delegates the choice to the router, which selects a model for the request rather than requiring the caller to hardcode one. Streaming responses are relayed as they arrive. When a self-hosted deployment lacks a key for some long-tail provider, requests can fall through to the hosted endpoint, which handles routing and billing for models the operator does not want to manage keys for; that path requires an account, whereas the self-hosted path does not. Deployment is a docker compose file: the container prints a generated API key on startup and serves on a local port. The project positions itself between a client library that only wraps providers, a closed hosted aggregator, and a purely local model runner, by being a server you run yourself that still has a managed escape hatch. The edition published here is limited to a single workspace, with multi-tenant routing reserved for the hosted product. It suits small teams and product builders who want one stable internal endpoint in front of many model vendors without operating a full gateway platform.

https://github.com/Continuum-AI-Corp/OrcaRouter-Lite

llm-routerllm-gatewayopenai-compatibleself-hostedbyokpython

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.