Awesome Infra for AI › Model Serving Frameworks

golf-mcp/golf

⭐ 840 Python repository created 2025-02-24

Golf is a comprehensive Python framework designed to simplify the development and deployment of AI agent server applications. It facilitates the creation of "MCP servers," enabling developers to define an agent's capabilities including tools, prompts, and resources using standard Python files within a structured directory. The framework automates the discovery, parsing, and compilation of these components into a runnable server, significantly reducing boilerplate code and accelerating development cycles. Key features of Golf include enterprise-grade authentication mechanisms such as JWT, OAuth Server mode, and API key support, ensuring secure agent operations. It also integrates built-in utilities for streamlined LLM interactions and automatic telemetry via OpenTelemetry for monitoring and tracing. This allows developers to focus primarily on implementing an agent's logic while Golf handles essential server infrastructure, security, and observability concerns. The framework provides a CLI for project scaffolding, a development server, and clear configuration options for aspects like host, port, transport protocols, and telemetry settings. It supports structured organization of tools, resources, and prompts, deriving component IDs automatically from file paths. With Golf, teams can quickly build, deploy, and scale secure AI agent infrastructure, including robust authentication, debugging, and continuous monitoring capabilities for production-ready AI agents.

https://github.com/golf-mcp/golf

AI agentsagent runtimeAI agent infrastructuremodel servingobservabilityAI authenticationLLM interactionspromptsPython frameworkOpenTelemetryAPI gateways

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...