Awesome Infra for AI › LLM Observability & Tracing

alibaba/loongsuite-pilot

⭐ 197 TypeScript added to this list on 2026-08-17 repository created 2026-06-04

LoongSuite Pilot is a local telemetry collector for AI coding agents, published by Alibaba. Development teams typically run more than one agent, and each records its activity in a different private format; Pilot gives them one collector that discovers those agents on a developer machine, deploys the required hooks or plugins, reads the resulting logs, sessions or data files, and normalizes everything into a shared GenAI event schema. Supported agents include Claude Code, Codex, Cursor and Cursor CLI, DeepSeek Harness, Hermes Agent, Kiro CLI and MiMo Code, each integrated through a hook, a plugin injection or session polling, and most reporting traces, logs, token usage and full conversation and tool-call detail. Normalized events can be exported to several destinations at once: JSONL files on disk, Alibaba Cloud SLS, a generic HTTP endpoint, or any OTLP trace backend, which lets the data land in whatever observability stack the team already runs. Because agent transcripts contain prompts, tool arguments and sometimes secrets, Pilot carries per-agent content-capture policy and secret masking that are applied before anything is exported. It also ships local operations — service status, restart, rollback of installed hooks — and a built-in dashboard that shows multi-agent token usage, sessions, requests, tools, models, providers and repository activity at a glance. Written in TypeScript and run as a local service, it is aimed at platform and engineering-productivity teams that need to answer practical questions about agent adoption: which agents are actually used, what they cost in tokens, what they touched, and whether their traffic is safe to forward to a central backend for analysis and audit.

https://github.com/alibaba/loongsuite-pilot

observabilitytelemetryopentelemetrytracingcoding-agentscollectorllmops

Also in LLM Observability & Tracing

langfuse/langfuse

Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.

comet-ml/opik

Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.

raga-ai-hub/RagaAI-Catalyst

RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.

Arize-ai/phoenix

Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...

VoltAgent/voltagent

VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...

traceloop/openllmetry

OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.

Helicone/helicone

Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.

Agenta-AI/agenta

Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.