langfuse/langfuse
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Awesome Infra for AI › LLM Observability & Tracing
Codex Usage Tracker is a local-first tool designed for developers utilizing OpenAI's Codex, providing detailed insights into token usage, credits, costs, and caching patterns. It reads local JSONL logs generated by Codex and offers deterministic analysis through MCP (Model Context Protocol) tools, a CLI, and an optional Evidence Console dashboard. The core purpose of this project is to allow developers to understand and optimize their Codex consumption without uploading sensitive logs to external services. It emphasizes privacy by processing all data locally and provides features for identifying high-cost threads, inefficient caching, and potential token waste. Users can interact with their usage data conversationally via a companion Codex skill, asking questions like "What drove my usage this week?" or "Find high-context, low-cache calls and link the exact supporting evidence." The tool facilitates in-depth analysis through its dashboard which presents readiness, freshness, bounded findings, and recent evidence. It helps uncover instances of token waste and suggests remediations by pointing to specific calls, threads, or findings. This allows developers to make informed decisions about their prompt engineering and model interactions, ultimately leading to more cost-effective and efficient use of AI models. Features include aggregate counters, an event index, and the ability to deduplicate cloned tasks while preserving provenance. It's built for individual developers seeking transparency and control over their local AI development costs and efficiency.
https://github.com/douglasmonsky/codex-usage-tracker
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.
Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.
RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.
Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...
VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...
OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.
Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.
Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.