Awesome Infra for AI › Model Serving Frameworks

superlinked/sie

⭐ 3360 Python repository created 2023-11-07

Superlinked Inference Engine (SIE) is an open-source inference server designed to streamline the serving of AI models for common tasks such as generating embeddings, reranking search results, and performing entity extraction. It offers a unified API for over 85 pre-configured models, covering dense, sparse, multi-vector, vision, and cross-encoder architectures, all quality-verified against MTEB benchmarks. SIE aims to replace fragmented model serving solutions with a single, comprehensive system that supports multiple models simultaneously with on-demand loading and LRU eviction. The project includes a full production stack, comprising a load-balancing gateway, KEDA autoscaling for efficient resource management (including scale-to-zero capabilities), Grafana dashboards for monitoring, and Terraform modules for deployment on GKE and EKS. It integrates with popular AI frameworks and tools such as LangChain, LlamaIndex, Haystack, DSPy, CrewAI, Chroma, Qdrant, and Weaviate, and provides an OpenAI-compatible /v1/embeddings endpoint for easy migration. SIE is packaged as a Docker container, making it easy to deploy from a laptop to a production Kubernetes cluster, and offers SDKs for Python and TypeScript.

https://github.com/superlinked/sie

inference serverembeddingsrerankingextractionLLMdeep learningNLPMLOpsKubernetesmodel servingproductionAI deploymentGPU inferencemodel optimization

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...