langfuse/langfuse
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Awesome Infra for AI › LLM Observability & Tracing
ax records what AI coding agent sessions actually did and feeds the result back into later sessions. A stop hook fires when a session ends, including sub-agent sessions, and asks the agent for a structured retrospective covering what was tried, what worked, what failed and what to do next. That retrospective is indexed as a typed experiment in a local graph alongside transcripts, git history and tool calls, so the knowledge a sub-agent accumulated does not disappear when the sub-agent does. From the accumulated sessions the tool identifies repeated friction, such as a command that fails several times before the right one is found, and turns each pattern into a small repository-specific proposal that a person triages one at a time. Accepted proposals become tracked experiments with verdicts recorded at seven, thirty and ninety days, which makes it possible to say whether a change actually helped rather than assuming it did. The same index answers operational questions about agent usage: which experiments are still open, which skills earned their keep, and what a branch cost in tokens. Everything runs locally against a DuckDB cache on the developer machine. There is no daemon and no server, ingest performs no outbound calls, and the only paths that leave the machine are explicit opt-in commands that publish aggregate counts rather than transcript content or code. Claude Code and Codex histories are supported, and setup can be driven either from the command line or by handing an onboarding prompt to the agent itself.
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.
RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.
Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...
VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...
OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.
Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.
Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.