langfuse/langfuse
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Awesome Infra for AI › LLM Observability & Tracing
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.
RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.
Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...
VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...
OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.
Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.
Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.
Latitude is an open-source AI monitoring platform that provides issue detection, human-aligned evaluations, and agent-native tracing for LLM applications and AI agents.
Pydantic Logfire is an observability platform for Python applications, providing detailed insights into production systems, including those leveraging LLMs and FastAPI, built on OpenTelemetry.
Laminar is an open-source observability platform purpose-built for AI agents, offering tracing, evaluation, AI monitoring, SQL access, dashboards, and data annotation for LLM-based applications.
A local proxy and trace viewer for AI coding agents, capturing and inspecting API traffic to debug agent behavior and analyze prompts, messages, and tool definitions.
OpenLIT is an open-source platform offering OpenTelemetry-native observability for LLMs, including GPU monitoring, guardrails, evaluations, prompt management, and API key vault, to streamline AI de...
A real-time desktop widget to monitor token usage, AI limits, and costs across various AI coding tools, featuring multi-device synchronization and historical usage trends.
Ragbits is a toolkit for rapid development and operation of GenAI applications, providing building blocks for LLM integration, RAG processing, multi-agent workflows, observability, and testing.
OpenInference provides conventions and instrumentation for OpenTelemetry to enable detailed tracing and observability of AI applications, especially those built with LLMs and agents.
Langtrace is an open-source, OpenTelemetry-based observability tool providing real-time tracing, evaluations, and metrics for LLM applications, including LLMs, LLM frameworks, and vector databases.
vLLora is a lightweight, real-time debugging and observability tool for AI agents, providing tracing and analysis of LLM interactions via an OpenAI-compatible API.
Local-first console that aggregates token usage, cache activity, model choice and estimated cost across Claude Code and Codex sessions on every connected machine.
TraceRoot is an open-source observability and self-healing platform for AI agents, providing tracing, AI-powered debugging, and detectors for production issues like hallucinations and tool failures.
PandaProbe is an open-source agent engineering platform for collaboratively tracing, evaluating, monitoring, and debugging AI agents, with integrations for LangGraph, CrewAI, and other agent SDKs.
Agentacct is a local-first agent work intelligence tool for coding agents, providing a dashboard to track token usage, estimated costs, tasks, and work steps without external telemetry.
An official plugin for OpenClaw that exports agent traces to Opik for comprehensive LLM observability, monitoring, and evaluation.
agenttrail is a local, dependency-free observability board for AI coding agents that merges an agent-maintained plan with observed file changes into a live map of what the agent is doing right now.
Station is an open-source, Git-backed runtime for deploying and orchestrating intelligent multi-agent AI systems on self-hosted infrastructure with built-in evaluation and observability.
Real-time local observability dashboard for AI agent runtimes: auto-detects installed coding agents and meters their sessions, tools, models, providers and token usage in one view.
OpenLLMetry-JS provides open-source observability for LLM applications in JavaScript/TypeScript, built on OpenTelemetry to trace interactions with LLM providers and vector databases.
AgentPrism is an open-source library of React components for visualizing traces from AI agents, turning complex OpenTelemetry and Langfuse data into clear, debuggable diagrams.
TokenTelemetry is a 100% local, open-source observability dashboard for AI coding and autonomous agents, tracking token usage, costs, tool calls, and session traces.
An open-source, local-first observability dashboard for OpenClaw AI agents, providing cost analytics, live monitoring, deep session inspection, and security auditing.
Palico AI is an integrated framework for iterative development, evaluation, and production of LLM applications, offering tools for building, improving performance, and debugging.
Hermes Labyrinth is a read-only observability plugin for Hermes Agent, visualizing autonomous agent journeys and interactions with prompts, tools, and memory into a navigable map.
Records coding-agent sessions such as Claude Code or Codex through a local proxy, then replays them offline byte-for-byte or forks mid-run onto a different model to compare outcomes.
Opik MCP (Model Context Protocol) server for integrating AI hosts (Claude Code, Cursor, VS Code Copilot) with Opik for LLM observability, evaluation, and prompt management.
Open measurement standard and local tooling that audits AI-agent usage telemetry for double-counted retries and inflated metrics, and screens Codex session logs for repeated failures.
Alibaba's local telemetry collector for AI coding agents: discovers installed agents, installs hooks, normalizes activity into a shared GenAI schema and exports logs and traces to chosen backends.
KnowledgeOps Agent is an enterprise-grade Spring AI platform designed for multi-tenant RAG, tool calling, and agent workflow orchestration, featuring robust security, observability, and evaluation ...
Tracks and analyzes local Codex (OpenAI) token usage, costs, and thread patterns through a local-first dashboard and CLI, aiding in cost optimization and waste reduction for AI developers.
Agent Inspect provides local execution trees, tracing, and debugging for TypeScript AI agents, offering in-depth insights into agent runs without cloud dependencies.
Axon is an OpenTelemetry-native CLI for local LLM observability, providing real-time dashboards to monitor and debug LLM/agent traces without cloud accounts.
An open-source, modular LLMOps stack combining LiteLLM for LLM API unification, routing, and cost control, with Langfuse for detailed observability, prompt versioning, and performance evaluation in...
Desktop and hook-based service that captures Codex and Claude Code session activity and surfaces continuous-improvement recommendations in a personal dashboard.
Agentwatch is a platform-agnostic observability framework for monitoring AI agent interactions, LLM calls, and tool usage across various AI development frameworks, providing real-time insights and ...
TMA1 provides local-first observability for LLM agents, capturing every LLM call, routing insights back to the agent via hooks and MCP tools, and offering a dashboard for detailed analysis.
Local-first observability and memory layer for AI coding agents that indexes every session into a typed graph and turns recurring friction into reviewable, tracked experiments.
inferock-bench is a local LLM cost-tracking proxy that provides independent, per-call receipts for OpenAI, Anthropic, Gemini, and OpenRouter, helping users audit bills, track token usage, and ident...
Local-first dashboard that turns Claude, Codex, Cursor and other coding-agent session logs into one view of token usage, estimated cost and context pressure.
Passive agent observability tool that reconstructs full agent turns, tool calls and LLM interactions purely from network traffic, with zero SDK, proxy or code changes.