langfuse/langfuse
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Awesome Infra for AI › LLM Observability & Tracing
AgentMeasure is an open measurement layer for AI-agent telemetry, built around the claim that agent usage metrics often conflate execution facts with logical operations: a retry counted as two requests instead of one logical operation, a reasoning-token subset added into a total it already belongs to, or a cache hit counted as a new measurement. Its conformance pack runs as a CLI or GitHub Action against a telemetry fixture and reports each check as PASS, FAIL or UNPROVABLE, treating missing evidence as an explicit, disclosed result rather than a silently zeroed one. A companion tool, Healthcheck, reads existing Codex Desktop rollout logs locally, with no network calls, and produces a terminal summary and HTML report flagging duplicate records, retry chains and consecutive tool failures; Claude Code support and independent verification of Codex CLI figures are still in progress. Beyond conformance checks, an experiment engine (lab/) lets users preregister and run task-set x harness x factor experiments that compute effect sizes with confidence intervals along a Reach-Choice-Success-Consumption funnel, aimed at making agent capability comparisons honest rather than cherry-picked. The project positions itself as the measurement foundation the emerging outcome-based AI pricing market (Zendesk, Intercom Fin, Sierra) currently lacks: a shared definition of what counts as one resolution or one operation, published as a free, local-first, opt-in specification and toolset rather than a paid ranking or marketplace. It targets teams building or auditing AI-agent products and tool authors who publish usage or success-rate claims and want to verify them with reproducible evidence instead of self-reported numbers.
https://github.com/roy-tong/AgentMeasure
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.
RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.
Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...
VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...
OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.
Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.
Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.