vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
TensorSharp provides a dedicated .NET-based solution for inferencing GGUF-formatted large language models and text-diffusion models. Its core purpose is to enable efficient and platform-flexible deployment of these AI models across Windows, macOS, and Linux environments, leveraging GPU capabilities for optimized performance. The project includes a command-line interface for direct interaction, a web-based chatbot for user frontends, and OpenAI/Ollama-compatible HTTP APIs, making it suitable for programmatic integration into other applications. TensorSharp supports various compute backends including GGML Metal for Apple Silicon, GGML CUDA for NVIDIA GPUs, GGML Vulkan for vendor-neutral GPU acceleration, and CPU-only options, ensuring broad hardware compatibility. It explicitly focuses on the inference phase, addressing the operational aspects of serving AI models rather than their training or fine-tuning. The engine is designed to handle different quantization levels to balance performance and memory usage, supporting a range of GGUF models from various families.
https://github.com/zhongkaifu/TensorSharp
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...