Awesome Infra for AI › LLM Gateways & Proxies

intentee/paddler

⭐ 1677 Rust repository created 2024-04-27

Paddler is an open-source, self-contained platform designed for deploying, serving, and scaling Large Language Models (LLMs) and Vision-Language Models (VLMs) on private infrastructure. It functions as an LLM load balancer, primarily utilizing a built-in llama.cpp engine for inference, supporting both CPU and GPU operations. The platform aims to provide privacy, reliability, cost control, and independence from closed-source model providers, making it suitable for product teams needing LLM inference, DevOps/LLMOps teams scaling models, and organizations with strict compliance or privacy requirements. Key features include LLM-specific load balancing, dynamic agent integration with autoscaling capabilities, request buffering for scaling from zero, and dynamic model swapping. It offers a web admin panel for management, monitoring, and testing, as well as observability metrics. Paddler consists of two main components: a 'balancer' that distributes incoming requests and exposes inference, management, and web admin services; and 'agents' that generate tokens and embeddings through dedicated slots. The project also provides a desktop application for more casual local AI cluster setups. It is designed for straightforward installation as a single binary and focuses on providing an efficient, self-hosted solution for production LLM inference.

https://github.com/intentee/paddler

aillamacppllmllmopsload-balancermodel servinginferenceenterprise llmself-hosted ai

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.