Awesome Infra for AI › Weekly › 2026-08-24

2026-08-24

16 projects added

AI Governance & Compliance

Floe-Labs/floe-guard

In-process spend meter and budget gate for AI agents that hard-stops the next model call, voice turn or paid tool invocation before it crosses a configured cost ceiling.

AI Safety & Guardrails

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

decionis/agent-safe-pipeline

Reference architecture and TypeScript library that routes AI agent actions through an independent authorization boundary, turning proposals into policy verdicts, human approvals and single-use grants.

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.

Inference Optimization

wangeai/K3Flight

Single-file Linux inference server that runs a 2.8 trillion parameter Kimi K3 checkpoint on CPU by streaming weights from storage, keeping the measured resident working set near 55 GB.

ARahim3/mlx-dspark

Native MLX port of the DSpark and DFlash speculative decoding drafters, giving lossless multi-fold faster LLM decoding on Apple Silicon with an OpenAI-compatible serving mode.

LLM Evaluation & Testing

Jwuthri/Tracely-ai

Trace-native CI/CD for AI agents that grades production traces, clusters failures, freezes bad runs into hermetic replayable regression cases and blocks the pull request that would ship them again.

LLM Gateways & Proxies

Continuum-AI-Corp/OrcaRouter-Lite

Self-hosted OpenAI-compatible LLM router that fronts more than a hundred models with bring-your-own-key access, streaming, automatic model selection and an optional hosted fallback.

HarnessRouter/harnessrouter

Self-hosted unified API in front of agent harnesses such as Codex, Claude Code and Hermes, implementing the Unified Harness Protocol with sessions, streaming, files and cancellation.

askalf/dario

Local OpenAI- and Anthropic-compatible proxy that exposes a Claude subscription to other AI tools, with session-affinity routing, multi-account pooling and drift detection for long agent runs.

azrtydxb/Fastllm-proxy

Rust LLM router presenting one OpenAI-compatible endpoint in front of eighty providers and local vLLM or SGLang backends, with cache-affinity routing, RBAC, budgets and no I/O on the request path.

LLM Observability & Tracing

Necmttn/ax

Local-first observability and memory layer for AI coding agents that indexes every session into a typed graph and turns recurring friction into reviewable, tracked experiments.

Prompt Management

newdee/prompt-shelf

Self-hosted prompt management service that applies Git-like version control to prompts, exposing them over a REST API with caching, authentication and role-based access.

Vector Databases & Retrieval Infrastructure

doobidoo/mcp-memory-service

Self-hosted persistent memory backend for AI agents that stores decisions and context in a semantic store with a typed knowledge graph, served over MCP, REST, CLI and a dashboard.

ggozad/haiku.rag

Agentic retrieval-augmented generation library for local document collections, combining hybrid vector and full-text search, reranking and multimodal retrieval on an embedded LanceDB store.

DrDroidLab/open-index

Structured context layer for AI agents that builds a searchable, continuously refreshed graph of domain entities and relationships from schemas and connectors.

Newer issue Older issue