Awesome Infra for AI › Model Serving Frameworks

containers/podman-desktop-extension-ai-lab

⭐ 300 TypeScript repository created 2023-12-19

Podman AI Lab is an open-source extension for Podman Desktop that enables users to work with Large Language Models (LLMs) in a local containerized environment. It offers a recipe catalog for common AI use cases, a curated selection of open-source models, and a playground for learning, prototyping, and experimentation. The extension helps users integrate AI into their applications without relying on external infrastructure, ensuring data privacy and security. It utilizes Podman machines to run inference servers for LLM models and AI applications, supporting model formats like GGUF, Pytorch, and Tensorflow. Users can download pre-curated AI models, start them as model services (inference servers) exposed via a chat API, and interact with them in integrated Playground environments to test capabilities and optimize parameters. The platform also supports AI applications as interconnected containers, offering a Recipes Catalog with detailed explanations and sample applications for use cases like chatbots and code generators. Hardware requirements include a minimum of 12GB RAM and 4 CPUs, with GPU acceleration on the roadmap. The extension seamlessly integrates with Podman Desktop, allowing for easy installation and management of local AI workloads.

https://github.com/containers/podman-desktop-extension-ai-lab

aicontainersinference-serverllmslocalpodmanllm-opsmodel-servingexperimentation

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...