Awesome Infra for AI › Model Serving Frameworks

Tejas-TA/predikit

⭐ 392 Python repository created 2026-05-25

Predikit is a Python library designed to transform trained scikit-learn or XGBoost models into LLM-callable tools, making it seamless to integrate traditional machine learning into AI agent workflows. It achieves this by automatically generating JSON schemas compatible with OpenAI's function calling API, handling typed inputs and outputs, and eliminating boilerplate code. Developers can define Pydantic schemas for model inputs, and Predikit wraps the model into a `ModelTool` object that can be directly invoked, converted to an OpenAI function schema, or integrated into LangChain. The library ensures input validation and provides clear error messages for schema mismatches. It also supports `ToolRegistry` for grouping multiple tools, and `ModelEnsemble` for running and reconciling outputs from multiple models, with strategies like `collect`, `mean`, `vote`, and weighted variations. Predikit is particularly useful for operationalizing existing ML assets within new LLM-powered applications, enabling AI agents to leverage a wider range of predictive capabilities with minimal effort.

https://github.com/Tejas-TA/predikit

model servingLLMAI agentsscikit-learnXGBoostOpenAI function callingLangChainPydanticmodel integrationinferenceML ops

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...