Awesome Infra for AI › Inference Optimization

ome-projects/ome

⭐ 514 Go repository created 2025-05-19

OME (Open Model Engine) is a sophisticated Kubernetes operator specifically engineered for orchestrating the serving of Large Language Models (LLMs) in production environments. It automates critical aspects of LLM operations, including model lifecycle management, intelligent runtime selection, and efficient resource allocation. The operator treats models as first-class custom resources, capable of parsing model metadata, supporting distributed storage with encryption, and handling various model formats like SafeTensors, PyTorch, and TensorRT. OME integrates with leading inference engines such as SGLang and vLLM for high-throughput serving and supports optimized deployment patterns including prefill-decode disaggregation and multi-node inference. It features specialized GPU bin-packing scheduling for maximizing cluster efficiency, dynamic re-optimization, and hardware-aware scheduling through AcceleratorClass resources. Deep integration with Kubernetes components like Kueue for gang scheduling, LeaderWorkerSet for resilient deployments, and KEDA for autoscaling ensures robust and scalable LLM inference. OME also provides a web console for management and includes built-in benchmarking capabilities for performance evaluation. Its focus is purely on the operational aspects of serving already-trained LLMs within a Kubernetes ecosystem, distinguishing it from general-purpose MLOps platforms or training-focused tools.

https://github.com/ome-projects/ome

kubernetesllm-servinggpu-schedulingmodel-managementinference-optimizationsglangvllmtritonkubernetes-operatormodel-lifecycle

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.