vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Rapid-MLX is an open-source (Apache 2.0) LLM inference server for Apple Silicon Macs, built on Apple's MLX framework. It exposes OpenAI- and Anthropic-compatible HTTP APIs so existing clients and coding agents (Claude Code, Codex, and similar) can point at it without code changes, and ships both a command-line server and a native macOS app. The project's stated focus is reliable tool calling: it aims to reproduce the tool-call behavior coding agents expect from hosted APIs, rather than only optimizing raw token throughput. The authors publish benchmark comparisons against Ollama, reporting roughly 3x aggregate decode throughput at 8 concurrent streams on a Qwen3.6-35B-A3B model measured on an M2 Pro, along with the method and raw data behind the number and a disclosed list of cases where it is slower. It is distributed via PyPI and Homebrew and requires Python 3.10+ and an Apple Silicon Mac (M1 or later). The project is under active development, with CI, a Discord community, and a companion DeepWiki reference. Rapid-MLX targets developers who want to run open-weight LLMs locally on Mac hardware for use with coding agents and tool-calling workflows, as an alternative to cloud-hosted inference or other local runtimes such as Ollama or llama.cpp.
https://github.com/raullenchai/Rapid-MLX
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...