langfuse/langfuse
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Awesome Infra for AI › LLM Observability & Tracing
Arize AI's Phoenix is an open-source AI observability platform specifically designed for large language model (LLM) applications. It enables users to experiment with, evaluate, and troubleshoot their AI systems effectively. The platform offers comprehensive tracing capabilities for LLM application runtimes, utilizing OpenTelemetry-based instrumentation to provide deep insights into execution flows. For evaluation, Phoenix allows benchmarking application performance using LLM-based response and retrieval evaluations, facilitating robust testing before deployment. It also supports the creation and versioning of datasets for experimentation, evaluation, and fine-tuning purposes, ensuring data integrity and reproducibility. Users can track and evaluate changes to prompts, LLMs, and retrieval mechanisms through dedicated experiment management features. The integrated playground allows for optimizing prompts, comparing different models, adjusting parameters, and replaying traced LLM calls to refine application behavior. Furthermore, Phoenix offers advanced prompt management functionalities, including version control, tagging, and experimentation to systematically manage and test prompt changes. An opt-in, permission-gated agent, PXI, is built into the product to assist with debugging traces, iterating on prompts, and navigating the platform. Phoenix is built to be vendor and language agnostic, providing out-of-the-box support for popular frameworks like OpenAI Agents SDK, Claude Agent SDK, LangChain, LlamaIndex, and various LLM providers such as OpenAI, Anthropic, Google GenAI, and AWS Bedrock.
https://github.com/Arize-ai/phoenix
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.
RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.
VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...
OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.
Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.
Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.
Latitude is an open-source AI monitoring platform that provides issue detection, human-aligned evaluations, and agent-native tracing for LLM applications and AI agents.