Awesome Infra for AI › Model Serving Frameworks

basetenlabs/truss

⭐ 1209 Python repository created 2022-07-06

Truss is an open-source command-line interface (CLI) tool designed to simplify the packaging, deployment, and serving of AI/ML models in production environments. It allows developers to define their model's serving logic in Python alongside weights and dependencies, abstracting away the complexities of containerization, Docker, Kubernetes configuration, and GPU setup. Truss supports a wide array of Python frameworks, including popular ones like `transformers`, `diffusers`, PyTorch, TensorFlow, as well as optimized inference engines such as vLLM, SGLang, and TensorRT-LLM. While it integrates seamlessly with Baseten for deployment, it also supports deployment to self-managed infrastructure. Key features include a fast development loop with live reload, built-in support for GPUs, secrets, caching, and autoscaling, making it suitable for production-ready inference. The tool focuses on enabling a "write once, run anywhere" approach, ensuring consistent model behavior from development to production. Its primary purpose is to streamline the operational aspects of serving trained AI/ML models, emphasizing ease of use and production readiness for inference workloads.

https://github.com/basetenlabs/truss

artificial-intelligencemachine-learningmodel-servinginference-servermodel-deploymentMLOpsopen-sourcepackagingGPULLM

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...