langfuse/langfuse
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Awesome Infra for AI › LLM Observability & Tracing
Penelopa.ai is a monitoring and recommendation tool aimed at people who use AI coding agents such as Codex and Claude Code daily. A platform-specific installer script downloads a private Node/npm runtime, configures hooks inside the coding agent's settings (Stop and SessionEnd events in Codex, an equivalent mechanism in Claude Code), and builds a desktop companion app on macOS or Windows; on Linux only the hooks are supported. Once connected, captured session events, including messages and tool input/output, are uploaded to a hosted dashboard where users can review activity, per-session reports and practical recommendations for improving how they work with their coding agent. The desktop app also runs local diagnostics and self-repair commands, keeps working in the background via a system tray icon after the window is closed so queued events still get delivered, and stores credentials either in a local hook config file or, for the desktop app, in the OS-native credential store such as Keychain or DPAPI. Installation explicitly requires the user to review and trust the generated hooks before they take effect, and offers uninstall and data-purge options. Because analysis and delivery rely on a hosted backend, most of the functional value depends on that service rather than being purely local or open. Penelopa.ai targets individual developers and teams who want ongoing, low-effort visibility into how their AI coding agents are behaving and where their usage could improve, rather than teams building or auditing generic agent telemetry infrastructure themselves.
https://github.com/chigwell/Penelopa.ai
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.
RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.
Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...
VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...
OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.
Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.
Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.