Awesome Infra for AI › LLM Gateways & Proxies

wink-run/tokenbank

⭐ 105 JavaScript added to this list on 2026-09-28 repository created 2026-05-09

Token Bank is a local gateway and desktop/CLI application that centralizes access to multiple AI coding agents and their upstream model providers. It runs a local HTTP endpoint that implements Anthropic Messages, OpenAI Chat, and Codex Responses protocol adapters, so agents such as Claude Code, Codex CLI, Cursor, and OpenCode can be pointed at it, via a CLI environment-variable shim or a one-click config-file patch, without modifying the agent itself. Once agents are onboarded, the gateway provides full-chain usage tracing across devices, consolidating subscription, pay-as-you-go, free-tier, and locally-hosted-model usage (for example through Ollama) into one accounting view, and a routing layer that can automatically select among these options by learned usage patterns or by declared task type such as design, repo-qa, chore, or debug, along with optional lossless context compression. It also includes a built-in MCP relay that projects a user's accumulated MCP servers, skills, and prompts onto whichever agent is active, and a personalization layer that adapts recommendations from observed usage. A distinguishing feature beyond typical LLM gateways is an opt-in peer-to-peer compute-sharing layer, through which idle local or subscription capacity can be contributed to a community network and earn credits, and remote community agents can run on a contributor's machine. It ships as an Electron desktop app for macOS and Windows as well as a CLI and a Docker-based web UI. The project targets individual developers running several AI coding tools who want unified usage visibility, cost-aware routing, and shared-resource tooling in one local layer.

https://github.com/wink-run/tokenbank

llm-gatewayai-coding-agentsroutingcost-trackingmcplocal-first

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.