Awesome Infra for AI › Model Serving Frameworks

Zyora-Dev/zse

⭐ 234 Python added to this list on 2026-08-17 repository created 2026-02-24

ZSE is a production LLM inference engine built without the usual machine-learning dependency stack: no PyTorch, no Triton, no bitsandbytes, no transformers. It is pure Python plus ctypes over a kernel compiler that emits CUDA, ROCm HIP and Metal code directly, which makes the installed package a few megabytes instead of a few gigabytes and removes the runtime deserialization and allocator behaviour that dominate startup elsewhere. The two headline properties are cold start and memory. Models are stored in a pre-quantized .zse format that is memory-mapped rather than deserialized, so a 7B model reported at seven seconds to serve on a T4 against roughly three and a half minutes for a comparable vLLM AWQ configuration, with similar ratios measured on L4, A10G, A100 and MI300X. VRAM figures follow the same pattern, since ZSE allocates what the model and cache need instead of claiming the device, allowing a 32B INT4 model to be served in about 22 GB on an MI300X. Single-sequence throughput is competitive on data-center parts and lags on smaller cards, which the project reports openly in its own comparison tables. The KV cache uses adaptive block sizes with token-level eviction rather than fixed blocks with LRU, model conversion is a one-time offline step, and retrieval-augmented generation is built in with hybrid retrieval and cross-encoder reranking. Installation is a single pip package followed by a serve command. Backends cover CUDA, ROCm and Metal, benchmarks were run on rented cloud GPUs and an Apple M1, the test suite is extensive, and the project is Apache 2.0 licensed and written for Python 3.11 and later.

https://github.com/Zyora-Dev/zse

inference-engineservingquantizationcudarocmmetalcold-startzero-dependency

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...