vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Solo CLI is a command-line interface designed for the rapid deployment and serving of AI models, with a particular focus on physical AI and robotics applications. It aims to streamline the process of getting AI models, including language, vision, and action models, to operate on physical hardware, specifically at the edge and on-device. The tool enables users to interact with these models directly from the terminal, offering capabilities for fine-tuning and serving models in real-world scenarios. Key features include model download from Solo Hub, authentication utilities, and comprehensive robotics operations such as motor setup, calibration, teleoperation, data recording, model training (e.g., ACT or SmolVLA policies), and inference. It also supports local AI deployment with various server types like Ollama, vLLM, and llama.cpp, providing OpenAI-compatible API references. The CLI assists in setting up the environment, checking running models, system status, and stopping services. It emphasizes context-aware intelligence and is specialized for mission-critical tasks where edge processing is crucial for efficient operation. Additionally, it offers integration with tools like OpenClaw for AI-driven automation of setup and calibration tasks.
https://github.com/GetSoloTech/solo-cli
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...