Awesome Infra for AI › Model Serving Frameworks

pegainfer-project/pegainfer

⭐ 716 Rust added to this list on 2026-08-10 repository created 2026-02-17

PegaInfer is an LLM inference server built from scratch in Rust and CUDA. It carries no Python framework runtime at all: every GPU kernel and every scheduler is hand-written, and the default build serves Qwen3-4B and Qwen3-8B without any Python present. Model support is gated behind cargo features — qwen3 is on by default, while Qwen3.5 (hybrid Gated DeltaNet plus full attention), DeepSeek-V2-Lite (MoE with expert parallelism across two GPUs) and Kimi-K2-Instruct (MLA plus MoE with Marlin INT4 on an eight-GPU expert-parallel path) each require rebuilding the server with the matching feature flag. Model type is auto-detected from the checkpoint config.json, so the server only needs a --model-path. The Qwen3.5 and GLM5.2 feature builds invoke Triton or TileLang at build time to generate ahead-of-time kernels, but neither needs Python at runtime. The HTTP surface is an OpenAI-compatible /v1/completions endpoint supporting max_tokens, temperature, top_k, top_p and SSE streaming; sampling and logprob coverage varies by model line. The project publishes a head-to-head benchmark against vLLM 0.22.1 on a single RTX 5090 with Qwen3-4B in BF16, driven by the same vllm bench serve client: roughly 771 MB resident memory against vLLM 3814 MB, and a three-second cold start to HTTP-ready against vLLM 70 seconds cold or 32.7 seconds with a warm torch.compile cache. Requirements are a CUDA-capable GPU, the CUDA Toolkit and an NVIDIA R545 driver or newer; release builds are mandatory because debug CUDA builds are unusably slow. Windows is supported alongside Linux. It suits teams that want a small, fast-starting single-process inference server without a Python model framework in the deployment.

https://github.com/pegainfer-project/pegainfer

inferencellmrustcudamodel-servingopenai-apigpu

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...