Awesome Infra for AI › LLM Observability & Tracing

Arize-ai/phoenix

⭐ 11708 Python repository created 2022-11-09

Arize AI's Phoenix is an open-source AI observability platform specifically designed for large language model (LLM) applications. It enables users to experiment with, evaluate, and troubleshoot their AI systems effectively. The platform offers comprehensive tracing capabilities for LLM application runtimes, utilizing OpenTelemetry-based instrumentation to provide deep insights into execution flows. For evaluation, Phoenix allows benchmarking application performance using LLM-based response and retrieval evaluations, facilitating robust testing before deployment. It also supports the creation and versioning of datasets for experimentation, evaluation, and fine-tuning purposes, ensuring data integrity and reproducibility. Users can track and evaluate changes to prompts, LLMs, and retrieval mechanisms through dedicated experiment management features. The integrated playground allows for optimizing prompts, comparing different models, adjusting parameters, and replaying traced LLM calls to refine application behavior. Furthermore, Phoenix offers advanced prompt management functionalities, including version control, tagging, and experimentation to systematically manage and test prompt changes. An opt-in, permission-gated agent, PXI, is built into the product to assist with debugging traces, iterating on prompts, and navigating the platform. Phoenix is built to be vendor and language agnostic, providing out-of-the-box support for popular frameworks like OpenAI Agents SDK, Claude Agent SDK, LangChain, LlamaIndex, and various LLM providers such as OpenAI, Anthropic, Google GenAI, and AWS Bedrock.

https://github.com/Arize-ai/phoenix

LLM observabilityAI evaluationLLM tracingprompt managementAI agentsLLM experimentationOpenTelemetryML monitoringdebuggingMLOpsLLM testingfine-tuning

Also in LLM Observability & Tracing

langfuse/langfuse

Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.

comet-ml/opik

Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.

raga-ai-hub/RagaAI-Catalyst

RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.

VoltAgent/voltagent

VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...

traceloop/openllmetry

OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.

Helicone/helicone

Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.

Agenta-AI/agenta

Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.

latitude-dev/latitude-llm

Latitude is an open-source AI monitoring platform that provides issue detection, human-aligned evaluations, and agent-native tracing for LLM applications and AI agents.