Awesome Infra for AI › Model Serving Frameworks

avifenesh/memra

⭐ 2 Rust added to this list on 2026-08-17 repository created 2026-07-05

memra is an LLM inference engine written in Rust with CUDA kernels, serving an OpenAI-compatible HTTP API. Its organizing idea is per-device tuning: rather than picking settings that are adequate everywhere, a mechanism that wins on an RTX PRO 6000 but loses on an RTX 5090 becomes a default keyed on the detected device, so running the server with no flags gets that particular card's measured best. The primary targets are the RTX PRO 6000 Blackwell and the RTX 5090, with a compile-gated H100 lane that is explicitly secondary and sets no defaults. Correctness is treated as a gate on every speed feature: speculative decoding, CUDA graph capture and batched serving are each verified byte-identical to plain decode on a per-request basis, so throughput is not bought with silent output drift. safetensors is the primary and tuned model format going forward, with GGUF still supported as the path most checkpoints take. Deployment shape is one model per GPU, replicas across cards behind an admission proxy, and two-stage pipeline parallelism when a model does not fit on a single card, validated across four GPUs; tensor parallelism, peer-to-peer transfer and three-stage pipelines are in progress. A release installer selects the right prebuilt for the detected architecture, verifies the checksum, and installs the server together with generation, speculative-decoding and kernel-check binaries; prebuilts require Linux x86_64, a recent glibc, driver 580 or newer and the CUDA runtime. Model support is tracked per model, quantization and drafter combination rather than by format alone, and unsupported checkpoints are requested through issues. MIT licensed.

https://github.com/avifenesh/memra

inference-engineservingcudarustblackwellspeculative-decodingopenai-compatible

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...