vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
MUXI provides open-source infrastructure designed for deploying and operating AI agents in production environments. Unlike traditional frameworks or libraries, MUXI functions as a dedicated server where agents are treated as native primitives. It offers integrated features for orchestration, multi-tier memory management (buffer, persistent, and vector), comprehensive observability with a wide range of event types and export options, and large-scale deployment capabilities. The platform supports declarative agent definitions through '.afs' files, allowing for version-controlled and auto-discovered AI systems. It comes with built-in multi-tenancy, RBAC, and OAuth, ensuring isolation and secure access. MUXI also integrates with over 1,000 tools via the MCP (Model-Controller-Programmer) and supports more than 21 LLM providers, offering flexibility and preventing vendor lock-in. It emphasizes self-hostability for data privacy and control, providing a robust solution for platform builders, internal tool developers, and any organization looking to deploy intelligent agents without reinventing infrastructure components. MUXI aims to simplify the deployment of stateful, API-accessible agents, making complex AI agent operations manageable through simple commands and declarative configurations.
https://github.com/muxi-ai/muxi
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...