Awesome Infra for AI › Model Serving Frameworks

bentoml/OpenLLM

⭐ 12552 Python repository created 2023-04-19

OpenLLM is a flexible framework that enables developers to easily deploy and serve open-source Large Language Models (LLMs) and custom models. It provides the capability to wrap these models into an OpenAI-compatible API endpoint with a single command, simplifying integration with existing tools and applications that leverage the OpenAI API specification. The platform includes a built-in chat UI for interaction and leverages state-of-the-art inference backends to optimize performance. OpenLLM supports a wide array of popular LLMs such as Llama, Mistral, and Qwen, and allows for the addition of custom model repositories. It integrates seamlessly with deployment tools like Docker and Kubernetes, and supports cloud deployment via BentoCloud, making it suitable for enterprise-grade AI infrastructure. A key feature is its ability to bundle models into 'Bentoml' for easy sharing and deployment across different environments. The project is focused on the serving and operational aspects of LLMs, providing a robust solution for MLOps engineers and developers looking to run their own LLM inference services efficiently.

https://github.com/bentoml/OpenLLM

bentomlllamallm-inferencellm-opsllm-servingllmopsmlopsmodel-inferenceopen-source-llmopenllm

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...

superduper-io/superduper

SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.