vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
SuperDuperDB is an innovative framework designed to bridge the gap between AI models and traditional databases, allowing developers to embed AI capabilities directly within their data stores. It supports various database backends like MongoDB, PostgreSQL, SQL databases, Redis, and Snowflake. The platform facilitates the deployment and serving of large language models (LLMs) and other AI models alongside your operational data, eliminating the need for standalone vector databases or complex ETL pipelines for AI inference. Key features include the ability to store, manage, and query outputs from any AI model, integrate models as database procedures, and perform real-time model inference directly on incoming data. It also supports Retrieval Augmented Generation (RAG) patterns by enabling semantic search and integrating vector embeddings within the database structure. SuperDuperDB aims to simplify the development and operation of AI applications, especially AI agents, by providing a unified environment for data, models, and inference, thereby streamlining MLOps for the inference side. It allows users to bring their own models or leverage pre-trained ones, enabling model serving and lightweight orchestration of AI workflows directly from the database.
https://github.com/superduper-io/superduper
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...