Awesome Infra for AI › Model Serving Frameworks

peonist-ai/halogen-flash-server

⭐ 840 Shell added to this list on 2026-09-14 repository created 2026-08-26

halogen-flash-server is a specialized inference server built for a single hardware and model pairing: the Qwen3.8-Flash-Next model family running on AMD's Strix Halo GPU. Rather than aiming for portability, every kernel is written specifically for this combination, which the project argues lets it beat general-purpose runtimes on speed without sacrificing precision. Benchmarked against other Strix Halo-targeted runtimes (EngramHalo.cpp, ROCmFP4, CIRU-IU4) on a 32K-token prompt with a 256-token answer, it reports roughly 4x faster end-to-end latency, driven mainly by faster prefill, while running at a higher effective precision (5.53 bits per weight) than the fastest competitor. Speculative decoding is used purely as a speed optimization: at temperature 0, output is verified to be byte-identical to serial greedy decoding on every release, using both the model's own draft head and prompt-lookup decoding from the request's own text. Since version 0.7.0 it can also load llama.cpp GGUF checkpoints of the same model directly, running them on the same custom kernels with the same speculative decoding and identity guarantees. The server exposes an OpenAI-compatible HTTP API, ships as a container image with Podman/Docker instructions, and documents configuration for cache modes, context and memory sizing, and an opt-in 1M-token context mode. It targets users who specifically run this model on this GPU and want the fastest possible serving stack, rather than teams needing a general-purpose, multi-model, multi-hardware inference server.

https://github.com/peonist-ai/halogen-flash-server

inference serverquantizationspeculative decodingAMD Strix HaloOpenAI-compatible API

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...