Awesome Infra for AI › LLM Observability & Tracing

Continuum-AI-Corp/OrcaReplay

⭐ 278 TypeScript added to this list on 2026-09-07 repository created 2026-08-29

OrcaReplay records an unmodified coding agent by inserting itself as a local proxy (via environment variables) between the agent and the model API. Because model APIs are stateless and each turn resends the full conversation, a proxy positioned there sees the entire agent loop: every request, streamed response, tool call, and tool result. It supplements this with four more capture layers to catch what the model protocol alone misses: a PATH shim recording exit codes, timing, and which stream a byte came from; a JSON-RPC tee for MCP traffic; a shadow git index that snapshots the workspace on every turn; and an opt-in TLS-intercept layer for agents with no configurable base URL, such as a Codex CLI signed in with a ChatGPT subscription. Everything is written to one trace file under `.orca/runs/`. The core commands are `orca record ` to capture a run unmodified, `orca replay last` to replay it from disk with the network blocked (no tokens spent, identical result every time), and `orca replay last --from N --model X` (or `orca compare` across several models) to resume from a checkpoint — a point where the conversation prefix and workspace state are fully known — and continue with a different model, isolating the model choice as the only variable. A companion command extracts and sanitizes the coding harness's own assembled system prompt per model. Built in TypeScript/Node by the team behind the OrcaRouter multi-provider model gateway, OrcaReplay targets developers and teams debugging why a coding agent (Claude Code, Codex, Agents SDK, AI SDK, and others) produced a bad result, or comparing how different models handle the exact same recorded task, without modifying the agent or wrapping it in an SDK.

https://github.com/Continuum-AI-Corp/OrcaReplay

agent replaysession recordingdebuggingproxymodel comparisontracingcoding agents

Also in LLM Observability & Tracing

langfuse/langfuse

Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications, offering observability, prompt management, and evaluation capabilities.

comet-ml/opik

Opik is an open-source platform for comprehensive observability, evaluation, and optimization of LLM applications, RAG systems, and agentic workflows.

raga-ai-hub/RagaAI-Catalyst

RagaAI Catalyst is a Python SDK for comprehensive observability, monitoring, and evaluation of AI agents and LLM applications, offering tracing, debugging, and advanced analytics.

Arize-ai/phoenix

Phoenix is an open-source AI observability platform for LLM application experimentation, evaluation, and troubleshooting, providing tracing, evaluation, dataset management, prompt management, and a...

VoltAgent/voltagent

VoltAgent is an end-to-end AI Agent Engineering Platform offering an open-source TypeScript framework for building intelligent agents and a VoltOps Console for observability, automation, deployment...

traceloop/openllmetry

OpenLLMetry provides open-source observability for LLM applications by extending OpenTelemetry to capture traces and metrics from LLM providers, vector databases, and AI frameworks.

Helicone/helicone

Helicone is an open-source LLM observability platform and AI gateway that provides monitoring, evaluation, prompt management, and intelligent routing for large language models.

Agenta-AI/agenta

Agenta is an open-source LLMOps platform designed to accelerate the development of reliable LLM applications, offering integrated prompt management, evaluation, and observability features.