Awesome Infra for AI › Model Serving Frameworks

Michael-A-Kuykendall/shimmy

⭐ 5918 Rust repository created 2025-08-28

Shimmy is a lightweight, single-binary inference server written entirely in Rust, designed to serve GGUF (GGML Universal File Format) large language models (LLMs) with 100% OpenAI-compatible API endpoints. This allows existing AI tools and applications to seamlessly integrate with locally run LLMs without modification. A key feature is its custom **Airframe engine**, a pure-Rust WebGPU (WGSL) transformer runtime that eliminates the need for C++ toolchains or complex compilation, ensuring high-quality, deterministic output on any GPU (NVIDIA, AMD, Intel, integrated) through WebGPU. Airframe automatically determines model specifications from GGUF metadata and supports extended context via YaRN RoPE scaling. Further enhancing efficiency, Shimmy includes **TurboShimmy INT4 KV**, an on-GPU KV-cache compression system. This innovation reduces KV cache VRAM consumption by approximately 7x without significant quality degradation, making it possible to run larger models or longer contexts on GPUs with limited memory (e.g., Llama-3.2-3B on 4GB GPUs or 7B models on 6GB GPUs). TurboShimmy operates entirely within WGSL compute shaders, converting 32-bit floats to 4-bit integers for the KV cache and dequantizing on-the-fly, ensuring no CPU roundtrips. Shimmy supports a wide range of GGUF models across various architectures and quantizations, read directly from model metadata for flexible deployment. The project emphasizes being 'free forever,' relying on community sponsorship for continuous development.

https://github.com/Michael-A-Kuykendall/shimmy

api-servercommand-line-tooldeveloper-toolsggufhuggingfacehuggingface-modelshuggingface-transformersinference-serverllamallamacppllm-inferencelocal-aimachine-learningollama-apiopenai-compatiblerustrust-cratetransformerswebgpuwebgpu-shaders

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

superduper-io/superduper

SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.