Awesome Infra for AI › Weekly › 2026-09-14

2026-09-14

11 projects added

Educational Resources

datawhalechina/zero-to-sglang

Hands-on course that builds a miniature SGLang inference engine from scratch, then walks through the real SGLang codebase's serving optimizations.

LLM Evaluation & Testing

jsdhwfmax/EvalForge

Portable evaluation-evidence and policy-gate tool that turns Ragas, promptfoo and DeepEval results into versioned, CI-enforceable pass/fail reports for RAG systems.

LLM Gateways & Proxies

mnfst/llm-gateway

Open-source LLM gateway that routes agent and app requests to 300+ models through one OpenAI-compatible endpoint, with cost tracking, fallback and self-healing.

theagentrouter/agent-router

Envoy-powered open source control plane for AI and agent traffic, giving one OpenAI-compatible API across model providers and self-hosted inference and MCP servers.

RelayPlane/proxy

Local proxy for AI agent traffic that prices every request in real time, rolls costs up per run and agent, and lets teams cap or kill runaway spend before it drains a budget.

0xzr/freellmpool

CLI, Python library and local proxy that pools free and keyless tiers from 22 LLM providers behind one OpenAI-compatible endpoint with automatic failover.

LLM Observability & Tracing

roy-tong/AgentMeasure

Open measurement standard and local tooling that audits AI-agent usage telemetry for double-counted retries and inflated metrics, and screens Codex session logs for repeated failures.

chigwell/Penelopa.ai

Desktop and hook-based service that captures Codex and Claude Code session activity and surfaces continuous-improvement recommendations in a personal dashboard.

Netis/heron

Passive agent observability tool that reconstructs full agent turns, tool calls and LLM interactions purely from network traffic, with zero SDK, proxy or code changes.

Model Serving Frameworks

peonist-ai/halogen-flash-server

Inference server hand-tuned for exactly one GPU and one model, serving Qwen3.8-Flash-Next on AMD Strix Halo through an OpenAI-compatible API with speculative decoding.

Workflow Orchestration for AI

hanyeol/model-compose

Declarative, docker-compose-style YAML tool that wires models, agents, RAG pipelines and MCP servers into deployable AI services across HTTP, WebSocket and MCP protocols.

Newer issue Older issue