Awesome Infra for AI › Model Serving Frameworks

bentoml/BentoDiffusion

⭐ 389 Python repository created 2023-06-12

BentoDiffusion is a collection of example projects built on BentoML, demonstrating how to deploy and serve different models from the Stable Diffusion family. Its core purpose is to facilitate the self-hosting and operationalization of powerful text-to-image and text-to-video diffusion models like SDXL Turbo, Stable Diffusion 3, and FLUX.1. Users can leverage these examples to run diffusion models locally, interact with them via HTTP clients (like cURL or Python), and deploy them to cloud environments like BentoCloud or other custom infrastructure using OCI-compliant images generated by BentoML. The project emphasizes the serving and inference aspects of these AI models, providing a practical framework for putting pre-trained diffusion models into production. It includes detailed instructions for setting up the environment, installing dependencies, running a BentoML service, and deploying to cloud platforms, making it a valuable resource for developers and MLOps engineers looking to integrate advanced generative AI capabilities into their applications.

https://github.com/bentoml/BentoDiffusion

diffusion-modelsmodel-servingbentomlkubernetesstable-diffusionimage-generationvideo-generationinferencedeployment

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...