vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...
SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.
LightLLM is a lightweight, scalable, and high-performance Python-based framework specifically designed for Large Language Model (LLM) inference and serving.
OpenAI- and Anthropic-compatible LLM inference server and native Mac app for Apple Silicon, built on MLX and optimized for reliable tool calling by coding agents.
LoRAX is an inference server designed to efficiently serve thousands of fine-tuned LLMs using LoRA adapters on a single GPU, optimizing cost, throughput, and latency.
Superlinked Inference Engine (SIE) is an open-source inference server that unifies the serving of over 85 pre-configured models for embeddings, reranking, and extraction, supporting seamless deploy...
RamaLama is an open-source tool that simplifies local AI model serving for inference using OCI containers, abstracting hardware complexities and enabling container-centric development for AI.
Chitu is a high-performance inference framework for large language models, designed for efficiency, flexibility, and availability across various hardware platforms.
Roboflow Inference is a self-hostable server for deploying and managing computer vision models and AI workflows on edge devices or other infrastructure, supporting various models and vision tasks.
Sonar is a high-performance, vLLM-based inference engine for serving HuggingFace-compatible LLMs at scale, featuring continuous batching, PagedAttention, and extensive quantization support.
A vLLM-style inference server for Apple Silicon built on MLX: continuous batching, paged and prefix KV cache, and both OpenAI and Anthropic APIs from a single Metal-backed process.
xLLM is a high-performance LLM inference engine optimized for diverse AI accelerators, particularly Chinese hardware, focusing on efficient, low-latency, and high-throughput model serving.
A multi-stage serving runtime from the SGLang project for omni, speech and text-to-speech models, exposing OpenAI-compatible audio and chat endpoints with streaming output.
Truss is a CLI tool and framework for packaging, deploying, and serving AI/ML models in production, handling containerization, dependency management, and GPU configuration.
Rust library for generating text, sparse, and image embeddings and reranking scores locally via ONNX Runtime, without a Python or GPU dependency.
Lanarky is a Python web framework built on FastAPI, specifically designed for creating LLM-powered microservices with native streaming support for HTTP and WebSockets.
Nanoflow is a high-performance, throughput-oriented serving framework for Large Language Models (LLMs) that utilizes intra-device parallelism and asynchronous CPU scheduling.
OpenVINO Model Server is a high-performance serving system for AI/ML models, optimized for OpenVINO and Intel architectures, offering efficient model inference via gRPC or REST APIs including OpenA...
LangCorn is an API server that facilitates serving LangChain LLM applications and agents using FastAPI, simplifying deployment and providing a robust inference solution.
Mosec is a high-performance, Rust and Python-based ML model serving framework that provides dynamic batching, pipelined stages, and CPU/GPU support for efficient online inference.
Pipeless is an open-source framework for building and deploying real-time computer vision applications, managing multimedia pipelines, model inference, and multi-stream processing.
Golf is a Python framework for building and deploying AI agent servers, providing infrastructure for authentication, observability, and managing tools, prompts, and resources.
Inference server hand-tuned for exactly one GPU and one model, serving Qwen3.8-Flash-Next on AMD Strix Halo through an OpenAI-compatible API with speculative decoding.
LLM inference engine written entirely in Rust and CUDA with no PyTorch or ONNX runtime, serving Qwen3 through trillion-parameter Kimi-K2 over an OpenAI-compatible HTTP API.
ServerlessLLM is a system for efficiently serving, multiplexing, and fine-tuning large language models on shared GPUs with ultra-fast model loading and an OpenAI-compatible API.
RAGLight is a modular Python framework designed for Retrieval-Augmented Generation (RAG), offering flexible integration with various LLMs, embeddings, and vector stores, and supporting agentic RAG ...
A FastAPI skeleton application designed to speed up the deployment and serving of machine learning models in production.
Pinferencia is a lightweight Python library for deploying machine learning models as inference servers with auto-generated GUI and REST APIs, supporting various ML frameworks and Kserve API.
TensorSharp is a native .NET inference engine specifically designed for serving GGUF large language models (LLMs) and DiffusionGemma-style text-diffusion models, offering console, web-based, and Op...
Predikit bridges traditional ML models (scikit-learn, XGBoost) with AI agents by automatically generating LLM-callable tools with typed I/O and OpenAI function schemas, simplifying model integration.
BentoDiffusion offers example projects for self-hosting and deploying various diffusion models using BentoML for image and video generation through text prompts.
A Podman Desktop extension for local LLM experimentation and application development using containerized inference servers and a recipe catalog.
Flama is a production framework for serving predictive and generative AI models as APIs, supporting OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native Model Context Protoc...
Fast CLI for deploying and serving AI models, especially for physical AI and robotics, optimized for edge and on-device operations.
A GPU-accelerated REST API for real-time object detection inference using YOLOv3 and YOLOv4 Darknet models, deployable via Docker or Docker Swarm.
A zero-dependency LLM inference server that owns its whole stack — no PyTorch or Triton — emitting CUDA, ROCm and Metal kernels directly for fast cold starts and a small memory footprint.
A CPU-based inference API for YOLOv3 and YOLOv4 object detection models, designed for easy deployment via Docker and Docker Swarm with RESTful endpoints.
MUXI is an open-source AI application server providing production infrastructure for deploying and operating AI agents with built-in orchestration, memory, observability, and scaling capabilities.
A Rust and CUDA inference engine with OpenAI-compatible serving, tuned per device class for Blackwell consumer and workstation GPUs rather than compromising across every card.