Awesome Infra for AI › Model Serving Frameworks

raullenchai/Rapid-MLX

⭐ 3898 Python added to this list on 2026-09-28 repository created 2026-02-25

Rapid-MLX is an open-source (Apache 2.0) LLM inference server for Apple Silicon Macs, built on Apple's MLX framework. It exposes OpenAI- and Anthropic-compatible HTTP APIs so existing clients and coding agents (Claude Code, Codex, and similar) can point at it without code changes, and ships both a command-line server and a native macOS app. The project's stated focus is reliable tool calling: it aims to reproduce the tool-call behavior coding agents expect from hosted APIs, rather than only optimizing raw token throughput. The authors publish benchmark comparisons against Ollama, reporting roughly 3x aggregate decode throughput at 8 concurrent streams on a Qwen3.6-35B-A3B model measured on an M2 Pro, along with the method and raw data behind the number and a disclosed list of cases where it is slower. It is distributed via PyPI and Homebrew and requires Python 3.10+ and an Apple Silicon Mac (M1 or later). The project is under active development, with CI, a Discord community, and a companion DeepWiki reference. Rapid-MLX targets developers who want to run open-weight LLMs locally on Mac hardware for use with coding agents and tool-calling workflows, as an alternative to cloud-hosted inference or other local runtimes such as Ollama or llama.cpp.

https://github.com/raullenchai/Rapid-MLX

llm-inferenceapple-siliconmlxmodel-servingtool-callinglocal-llm

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...