Awesome Infra for AI › LLM Observability & Tracing

cloudshipai/station

⭐ 430 Go repository created 2025-07-28

Station is an open-source platform designed for building, testing, and deploying intelligent multi-agent AI systems. It allows teams to orchestrate specialist agents under coordinating orchestrators, similar to how human teams collaborate. A core feature is its built-in evaluation capability, which uses LLM-as-judge methods to automatically test agents. The platform supports a Git-backed workflow, enabling version control of agents like code and facilitating GitOps practices. Users can deploy agents with a single command. Station provides full observability through Jaeger traces for every execution, offering deep insights into agent behavior. It is designed for self-hosting, giving users full control over their data and infrastructure. The platform supports multiple AI providers, including CloudShip AI, OpenAI, Google Gemini, and Anthropic. It facilitates integration with various MCP (Multi-Agent Communication Protocol) clients and editors, such as Claude Code CLI, OpenCode, and Cursor, through environment variables and configuration files. Station also includes features like 'Faker tools' for generating mock data for safe development and 'MCP templates' for managing credentials securely. The project provides a web UI for configuration and onboarding guides to help users get started with creating agents, using faker tools, developing multi-agent hierarchies, and inspecting run traces.

https://github.com/cloudshipai/station

multi-agent systemsagent orchestrationLLM evaluationAI observabilityGitOpsself-hosted AILLMOpsAI deploymentprompt management

Also in LLM Observability & Tracing

langfuse/langfuse

Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.

comet-ml/opik

Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.

raga-ai-hub/RagaAI-Catalyst

RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.

Arize-ai/phoenix

Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...

VoltAgent/voltagent

VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...

traceloop/openllmetry

OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.

Helicone/helicone

Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.

Agenta-AI/agenta

Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.