Awesome Infra for AI › Model Serving Frameworks

jundot/omlx

⭐ 22523 Python repository created 2026-02-13

oMLX is a specialized LLM inference server designed to maximize performance on Apple Silicon devices. Its core functionality revolves around efficient LLM serving through advanced techniques like continuous batching and a sophisticated tiered KV cache system. The KV cache leverages both hot (RAM) and cold (SSD) tiers, allowing frequently accessed context to remain in memory while less critical data is offloaded to disk, supporting persistent context across requests and even server restarts. This is particularly beneficial for conversational AI and coding assistants. The server supports a variety of models including text LLMs, vision-language models (VLMs), embedding models, and rerankers. It offers multi-model serving with intelligent management features such as LRU eviction, manual load/unload, model pinning, and per-model TTLs to optimize resource usage. An intuitive admin dashboard provides real-time monitoring, model management, and configurable per-model settings without requiring server restarts. The project provides flexible deployment options, including a macOS application, Homebrew installation, or source build, and integrates seamlessly with OpenAI-compatible clients. Its focus on local, optimized inference for Apple Silicon makes it a powerful tool for developers looking to run and manage AI models efficiently on their Macs.

https://github.com/jundot/omlx

apple-siliconinference-serverllmmacosmlxopenai-apicontinuous-batchingkv-cachingmodel-servingllm-ops

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...

superduper-io/superduper

SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.