Awesome Infra for AI › Model Serving Frameworks

Model Serving Frameworks

46 projects

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

⭐ 93184 Python

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

⭐ 22523 Python

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

⭐ 21865 Python

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

⭐ 12552 Python

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

⭐ 8877 Python

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

⭐ 7046 Python

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

⭐ 6075 Go

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...

⭐ 5918 Rust

superduper-io/superduper

SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.

⭐ 5328 Python

ModelTC/LightLLM

LightLLM is a lightweight, scalable, and high-performance Python-based framework specifically designed for Large Language Model (LLM) inference and serving.

⭐ 4305 Python

raullenchai/Rapid-MLX

OpenAI- and Anthropic-compatible LLM inference server and native Mac app for Apple Silicon, built on MLX and optimized for reliable tool calling by coding agents.

⭐ 3898 Python added 2026-09-28

predibase/lorax

LoRAX is an inference server designed to efficiently serve thousands of fine-tuned LLMs using LoRA adapters on a single GPU, optimizing cost, throughput, and latency.

⭐ 3835 Python

superlinked/sie

Superlinked Inference Engine (SIE) is an open-source inference server that unifies the serving of over 85 pre-configured models for embeddings, reranking, and extraction, supporting seamless deploy...

⭐ 3360 Python

containers/ramalama

RamaLama is an open-source tool that simplifies local AI model serving for inference using OCI containers, abstracting hardware complexities and enabling container-centric development for AI.

⭐ 3071 Python

thu-pacman/chitu

Chitu is a high-performance inference framework for large language models, designed for efficiency, flexibility, and availability across various hardware platforms.

⭐ 2987 Python

roboflow/inference

Roboflow Inference is a self-hostable server for deploying and managing computer vision models and AI workflows on edge devices or other infrastructure, supporting various models and vision tasks.

⭐ 2467 Python

dphnAI/sonar

Sonar is a high-performance, vLLM-based inference engine for serving HuggingFace-compatible LLMs at scale, featuring continuous batching, PagedAttention, and extensive quantization support.

⭐ 1869 C++ added 2026-07-06

waybarrios/vllm-mlx

A vLLM-style inference server for Apple Silicon built on MLX: continuous batching, paged and prefix KV cache, and both OpenAI and Anthropic APIs from a single Metal-backed process.

⭐ 1611 Python added 2026-08-17

xLLM-AI/xllm

xLLM is a high-performance LLM inference engine optimized for diverse AI accelerators, particularly Chinese hardware, focusing on efficient, low-latency, and high-throughput model serving.

⭐ 1589 C++ added 2026-07-13

sgl-project/sglang-omni

A multi-stage serving runtime from the SGLang project for omni, speech and text-to-speech models, exposing OpenAI-compatible audio and chat endpoints with streaming output.

⭐ 1315 Python added 2026-08-17

basetenlabs/truss

Truss is a CLI tool and framework for packaging, deploying, and serving AI/ML models in production, handling containerization, dependency management, and GPU configuration.

⭐ 1209 Python

anush008/fastembed-rs

Rust library for generating text, sparse, and image embeddings and reranking scores locally via ONNX Runtime, without a Python or GPU dependency.

⭐ 1030 Rust added 2026-09-21

ajndkr/lanarky

Lanarky is a Python web framework built on FastAPI, specifically designed for creating LLM-powered microservices with native streaming support for HTTP and WebSockets.

⭐ 989 Python

efeslab/Nanoflow

Nanoflow is a high-performance, throughput-oriented serving framework for Large Language Models (LLMs) that utilizes intra-device parallelism and asynchronous CPU scheduling.

⭐ 975 Jupyter Notebook

openvinotoolkit/model_server

OpenVINO Model Server is a high-performance serving system for AI/ML models, optimized for OpenVINO and Intel architectures, offering efficient model inference via gRPC or REST APIs including OpenA...

⭐ 941 C++

msoedov/langcorn

LangCorn is an API server that facilitates serving LangChain LLM applications and agents using FastAPI, simplifying deployment and providing a robust inference solution.

⭐ 941 Python

mosecorg/mosec

Mosec is a high-performance, Rust and Python-based ML model serving framework that provides dynamic batching, pipelined stages, and CPU/GPU support for efficient online inference.

⭐ 902 Python

pipeless-ai/pipeless

Pipeless is an open-source framework for building and deploying real-time computer vision applications, managing multimedia pipelines, model inference, and multi-stream processing.

⭐ 853 Rust

golf-mcp/golf

Golf is a Python framework for building and deploying AI agent servers, providing infrastructure for authentication, observability, and managing tools, prompts, and resources.

⭐ 840 Python

peonist-ai/halogen-flash-server

Inference server hand-tuned for exactly one GPU and one model, serving Qwen3.8-Flash-Next on AMD Strix Halo through an OpenAI-compatible API with speculative decoding.

⭐ 840 Shell added 2026-09-14

pegainfer-project/pegainfer

LLM inference engine written entirely in Rust and CUDA with no PyTorch or ONNX runtime, serving Qwen3 through trillion-parameter Kimi-K2 over an OpenAI-compatible HTTP API.

⭐ 716 Rust added 2026-08-10

ServerlessLLM/ServerlessLLM

ServerlessLLM is a system for efficiently serving, multiplexing, and fine-tuning large language models on shared GPUs with ultra-fast model loading and an OpenAI-compatible API.

⭐ 715 Python

Bessouat40/RAGLight

RAGLight is a modular Python framework designed for Retrieval-Augmented Generation (RAG), offering flexible integration with various LLMs, embeddings, and vector stores, and supporting agentic RAG ...

⭐ 672 Python

eightBEC/fastapi-ml-skeleton

A FastAPI skeleton application designed to speed up the deployment and serving of machine learning models in production.

⭐ 601 Python

underneathall/pinferencia

Pinferencia is a lightweight Python library for deploying machine learning models as inference servers with auto-generated GUI and REST APIs, supporting various ML frameworks and Kserve API.

⭐ 543 Python

zhongkaifu/TensorSharp

TensorSharp is a native .NET inference engine specifically designed for serving GGUF large language models (LLMs) and DiffusionGemma-style text-diffusion models, offering console, web-based, and Op...

⭐ 540 C# added 2026-07-13

Tejas-TA/predikit

Predikit bridges traditional ML models (scikit-learn, XGBoost) with AI agents by automatically generating LLM-callable tools with typed I/O and OpenAI function schemas, simplifying model integration.

⭐ 392 Python

bentoml/BentoDiffusion

BentoDiffusion offers example projects for self-hosting and deploying various diffusion models using BentoML for image and video generation through text prompts.

⭐ 389 Python

containers/podman-desktop-extension-ai-lab

A Podman Desktop extension for local LLM experimentation and application development using containerized inference servers and a recipe catalog.

⭐ 300 TypeScript

vortico/flama

Flama is a production framework for serving predictive and generative AI models as APIs, supporting OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native Model Context Protoc...

⭐ 299 Python added 2026-06-29

GetSoloTech/solo-cli

Fast CLI for deploying and serving AI models, especially for physical AI and robotics, optimized for edge and on-device operations.

⭐ 298 Python

BMW-InnovationLab/BMW-YOLOv4-Inference-API-GPU

A GPU-accelerated REST API for real-time object detection inference using YOLOv3 and YOLOv4 Darknet models, deployable via Docker or Docker Swarm.

⭐ 275 Python

Zyora-Dev/zse

A zero-dependency LLM inference server that owns its whole stack — no PyTorch or Triton — emitting CUDA, ROCm and Metal kernels directly for fast cold starts and a small memory footprint.

⭐ 234 Python added 2026-08-17

BMW-InnovationLab/BMW-YOLOv4-Inference-API-CPU

A CPU-based inference API for YOLOv3 and YOLOv4 object detection models, designed for easy deployment via Docker and Docker Swarm with RESTful endpoints.

⭐ 217 Python

muxi-ai/muxi

MUXI is an open-source AI application server providing production infrastructure for deploying and operating AI agents with built-in orchestration, memory, observability, and scaling capabilities.

⭐ 188 added 2026-07-13

avifenesh/memra

A Rust and CUDA inference engine with OpenAI-compatible serving, tuned per device class for Blackwell consumer and workstation GPUs rather than compromising across every card.

⭐ 2 Rust added 2026-08-17