Awesome Infra for AI › LLM Gateways & Proxies

NadirRouter/NadirClaw

⭐ 656 Python repository created 2026-02-11

NadirClaw is an open-source, self-hosted LLM router and cost optimization tool designed to significantly reduce API expenses for applications utilizing large language models. It acts as an OpenAI-compatible proxy, automatically classifying incoming prompts as "simple" or "complex" and routing them to appropriately priced models. Simple prompts, which often constitute the majority of requests in typical coding sessions (e.g., formatting, basic questions), are directed to cheaper, smaller models (like Gemini Flash or Haiku), while complex requests requiring advanced reasoning are sent to premium models (like Claude Sonnet or GPT-5.2) or local models like Ollama. This intelligent routing mechanism is capable of cutting AI API costs by 40-70%. The project emphasizes local operation, ensuring API keys remain on the user's machine, and offers features such as fallback chains for model resilience and built-in cost tracking. It integrates seamlessly with any OpenAI-compatible client, including popular tools like Claude Code, Cursor, and Codex. NadirClaw provides a ~10ms classification overhead and includes a "Context Optimize" feature to compact bloated context, further saving input tokens. While the open-source version uses a simpler binary centroid classifier, the Nadir Pro (hosted) version offers a more advanced trained classifier with a high AUROC score and additional enterprise features like a live dashboard, team billing, and SSO. The tool is deployable via `pip` or a simple install script, followed by an interactive setup wizard to configure providers and models. Benchmarks demonstrate its effectiveness in both cost savings and quality preservation, with a low catastrophic-downgrade rate.

https://github.com/NadirRouter/NadirClaw

aiai-cost-reductionai-routerclaude-codecodexcost-optimizationgeminillmllm-gatewayllm-proxyllm-routermodel-routingollamaopenaiopenai-proxyopenclawprompt-routingproxypythonself-hostedllm-opsllm-inferencecost-control

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.