datawhalechina/zero-to-sglang
Hands-on course that builds a miniature SGLang inference engine from scratch, then walks through the real SGLang codebase's serving optimizations.
Awesome Infra for AI › Weekly › 2026-09-14
Hands-on course that builds a miniature SGLang inference engine from scratch, then walks through the real SGLang codebase's serving optimizations.
Portable evaluation-evidence and policy-gate tool that turns Ragas, promptfoo and DeepEval results into versioned, CI-enforceable pass/fail reports for RAG systems.
Open-source LLM gateway that routes agent and app requests to 300+ models through one OpenAI-compatible endpoint, with cost tracking, fallback and self-healing.
Envoy-powered open source control plane for AI and agent traffic, giving one OpenAI-compatible API across model providers and self-hosted inference and MCP servers.
Local proxy for AI agent traffic that prices every request in real time, rolls costs up per run and agent, and lets teams cap or kill runaway spend before it drains a budget.
CLI, Python library and local proxy that pools free and keyless tiers from 22 LLM providers behind one OpenAI-compatible endpoint with automatic failover.
Open measurement standard and local tooling that audits AI-agent usage telemetry for double-counted retries and inflated metrics, and screens Codex session logs for repeated failures.
Desktop and hook-based service that captures Codex and Claude Code session activity and surfaces continuous-improvement recommendations in a personal dashboard.
Passive agent observability tool that reconstructs full agent turns, tool calls and LLM interactions purely from network traffic, with zero SDK, proxy or code changes.
Inference server hand-tuned for exactly one GPU and one model, serving Qwen3.8-Flash-Next on AMD Strix Halo through an OpenAI-compatible API with speculative decoding.
Declarative, docker-compose-style YAML tool that wires models, agents, RAG pipelines and MCP servers into deployable AI services across HTTP, WebSocket and MCP protocols.