Awesome Infra for AI › Model Serving Frameworks

GetSoloTech/solo-cli

⭐ 298 Python repository created 2024-06-30

Solo CLI is a command-line interface designed for the rapid deployment and serving of AI models, with a particular focus on physical AI and robotics applications. It aims to streamline the process of getting AI models, including language, vision, and action models, to operate on physical hardware, specifically at the edge and on-device. The tool enables users to interact with these models directly from the terminal, offering capabilities for fine-tuning and serving models in real-world scenarios. Key features include model download from Solo Hub, authentication utilities, and comprehensive robotics operations such as motor setup, calibration, teleoperation, data recording, model training (e.g., ACT or SmolVLA policies), and inference. It also supports local AI deployment with various server types like Ollama, vLLM, and llama.cpp, providing OpenAI-compatible API references. The CLI assists in setting up the environment, checking running models, system status, and stopping services. It emphasizes context-aware intelligence and is specialized for mission-critical tasks where edge processing is crucial for efficient operation. Additionally, it offers integration with tools like OpenClaw for AI-driven automation of setup and calibration tasks.

https://github.com/GetSoloTech/solo-cli

model-servinginferenceedge-aiphysical-airoboticson-device-aiclimodel-deploymentollamavllmllama.cpp

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...