Awesome Infra for AI › LLM Gateways & Proxies

Kenza-AI/sagify

⭐ 442 Python repository created 2018-03-03

Sagify provides a streamlined interface for managing machine learning workflows and LLM operations on AWS SageMaker. Its core offering is an LLM Gateway, which acts as a unified entry point for interacting with both proprietary models like OpenAI's GPT series and open-source LLMs such as Llama-2 and various embedding models. This gateway abstracts away the complexities of deployment and integration, allowing users to focus on leveraging LLMs for chat completions, image generation, and embeddings through a simplified API. Sagify supports deployment of these models as SageMaker endpoints, handling the underlying infrastructure setup. Users can configure and deploy services like chat completions, image creations, and embeddings, choosing from a variety of supported models. The tool also facilitates the local deployment of a FastAPI-based LLM Gateway, connecting to the deployed SageMaker endpoints or external APIs like OpenAI. It aims to reduce the operational overhead typically associated with bringing ML and LLM models into production, making them more accessible for developers and ML engineers by automating much of the AWS SageMaker integration and LLM service orchestration. While it interacts with the AWS SageMaker training capabilities, its primary value proposition lies in simplifying the serving and inference aspects of models, particularly LLMs, through its gateway module.

https://github.com/Kenza-AI/sagify

llmopsllm-inferencellm-gatewaymodel-servingsagemakeropenaiopen-source-llmgenerative-ailangchain

Also in LLM Gateways & Proxies

diegosouzapw/OmniRoute

OmniRoute is a free AI gateway that unifies access to over 170 AI providers, offering token compression, auto-fallback, and aggregating free tiers to provide billions of free tokens monthly.

Mintplex-Labs/anything-llm

AnythingLLM is an all-in-one local-first AI application for chatting with documents, managing AI agents, and integrating with various LLMs and vector databases, offering dynamic model routing and m...

BerriAI/litellm

LiteLLM is an open-source AI Gateway and Python SDK providing a unified interface to over 100 LLM providers, with features like cost tracking, guardrails, load balancing, and observability.

QuantumNous/new-api

new-api is a unified LLM gateway and AI asset management system that enables aggregation, distribution, and cross-conversion of various LLMs into OpenAI, Claude, or Gemini compatible formats, offer...

Kong/kong

Kong Gateway is a cloud-native API and AI gateway offering high performance, extensibility via plugins, and advanced AI traffic capabilities including multi-LLM support, semantic security, and cach...

decolua/9router

9Router is an AI router and token saver that connects various AI coding tools to over 40 AI providers, optimizing usage with auto-fallback, quota tracking, and token compression.

apache/apisix

Apache APISIX is a dynamic, real-time, high-performance API Gateway that can also function as an AI Gateway, providing AI proxying, load balancing for LLMs, and robust security for AI agents.

lidge-jun/opencodex

OpenCodex is a universal local proxy that enables OpenAI Codex, Claude Code, and Grok Build to utilize any LLM provider, offering advanced model routing, account pooling, and API key management.