vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
oMLX is a specialized LLM inference server designed to maximize performance on Apple Silicon devices. Its core functionality revolves around efficient LLM serving through advanced techniques like continuous batching and a sophisticated tiered KV cache system. The KV cache leverages both hot (RAM) and cold (SSD) tiers, allowing frequently accessed context to remain in memory while less critical data is offloaded to disk, supporting persistent context across requests and even server restarts. This is particularly beneficial for conversational AI and coding assistants. The server supports a variety of models including text LLMs, vision-language models (VLMs), embedding models, and rerankers. It offers multi-model serving with intelligent management features such as LRU eviction, manual load/unload, model pinning, and per-model TTLs to optimize resource usage. An intuitive admin dashboard provides real-time monitoring, model management, and configurable per-model settings without requiring server restarts. The project provides flexible deployment options, including a macOS application, Homebrew installation, or source build, and integrates seamlessly with OpenAI-compatible clients. Its focus on local, optimized inference for Apple Silicon makes it a powerful tool for developers looking to run and manage AI models efficiently on their Macs.
https://github.com/jundot/omlx
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...
SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.