Awesome Infra for AI › Model Serving Frameworks

containers/ramalama

⭐ 3071 Python repository created 2024-07-24

RamaLama is an open-source development tool designed to streamline the local serving and operationalization of AI models for inference, leveraging the familiarity and benefits of OCI containers. It allows engineers to apply container-centric development patterns to AI workloads, abstracting away the complexities of host system configuration. The tool automatically detects available GPUs, pulls appropriate accelerated container images (e.g., for CUDA, HIP, Intel, Asahi), and manages dependencies and hardware optimization within the containerized environment. RamaLama supports multiple AI model registries, including OCI Container Registries, treating models similar to how container images are managed by tools like Podman and Docker. It provides common container commands for interacting with AI models, prioritizing security by running models in rootless containers, isolating them from the host, and defaulting to no network access with temporary data removal upon exit. The project integrates with various hardware accelerators, including NVIDIA (CUDA), AMD (HIP), Apple Silicon (Asahi), Intel GPUs, and Huawei Ascend (CANN), ensuring broad compatibility. Users can interact with models via REST API or as a chatbot. Installation is flexible, covering macOS (via installer), Fedora (dnf), PyPI (pip), and Windows (via Docker Desktop or Podman Desktop with WSL2). RamaLama’s core purpose is to simplify the deployment and serving of pre-trained AI/ML models in a production-like inference environment, focusing on operational aspects rather than training or model development.

https://github.com/containers/ramalama

aicontainerscudahacktoberfesthipinference-serverintelllamacppllmpodmanvllmmodel-servinginferencellm-opsgpu-accelerationcontainerization

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...