vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Flama is a robust framework designed to streamline the deployment and serving of both predictive and generative AI models in production environments. It allows users to turn any trained model, whether from scikit-learn, TensorFlow, PyTorch, or large language models (LLMs), into a production-ready API with minimal effort. The core of Flama is powered by Rust, ensuring high performance for routing, JSON encoding, request parsing, and compression. Key features include the ability to package models into a portable `.flm` artifact, allowing for consistent API exposure regardless of the original framework. It supports serving LLMs with OpenAI, Anthropic, and Ollama-compatible endpoints, which means existing clients for these platforms can interact with Flama-served models without modifications. Each served LLM comes with a built-in streaming chat UI, readily available at a `/chat/` endpoint, supporting features like Markdown and LaTeX. Flama also offers native support for the Model Context Protocol (MCP), enabling AI agents to discover and invoke tools, resources, and prompts declared with simple decorators. This facilitates advanced agent interactions by automatically deriving JSON schemas from type hints. Beyond AI-specific functionalities, Flama provides a comprehensive toolkit for building general APIs, including resource management with SQLAlchemy, dependency injection, adaptable schemas (Pydantic, Typesystem, Marshmallow), auto-generated OpenAPI documentation, and features like pagination and background tasks. Installation is straightforward via PyPI, with optional extras for specific functionalities like schema validation, database integration, or generative AI serving using backends like vLLM or MLX.
https://github.com/vortico/flama
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...