kenryu42/cc-safety-net
A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.
Awesome Infra for AI › Weekly › 2026-08-17
A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.
A Go tool that fingerprints LLM serving infrastructure on network endpoints, identifying which of 60+ AI services — Ollama, vLLM, LiteLLM, TGI and others — is running behind a port.
Adaptive authorization and runtime guardrails for AI coding agents: an MCP proxy or host hook that turns every tool call into an auditable PASS, AUTH or BLOCK decision before it executes.
An open-ended benchmark harness that hands a coding agent a live Unreal Engine 5 editor over MCP and measures whether it can build a whole open-world game unaided.
A Rust-native AI gateway from the creators of Apache APISIX: one OpenAI-compatible API in front of every provider, with routing, failover, guardrails, caching and observability in a single binary.
Self-hosted, OpenAI-compatible LLM gateway that gives each project one /v1 endpoint and API key while the operator manages upstream accounts, routing chains and fallbacks in one place.
A self-hosted LLM gateway in one Go binary that speaks four wire protocols, translates between them, fails over across providers, rotates upstream keys and ships a multi-user admin console.
Real-time local observability dashboard for AI agent runtimes: auto-detects installed coding agents and meters their sessions, tools, models, providers and token usage in one view.
Alibaba's local telemetry collector for AI coding agents: discovers installed agents, installs hooks, normalizes activity into a shared GenAI schema and exports logs and traces to chosen backends.
A vLLM-style inference server for Apple Silicon built on MLX: continuous batching, paged and prefix KV cache, and both OpenAI and Anthropic APIs from a single Metal-backed process.
A multi-stage serving runtime from the SGLang project for omni, speech and text-to-speech models, exposing OpenAI-compatible audio and chat endpoints with streaming output.
A Rust and CUDA inference engine with OpenAI-compatible serving, tuned per device class for Blackwell consumer and workstation GPUs rather than compromising across every card.
A zero-dependency LLM inference server that owns its whole stack — no PyTorch or Triton — emitting CUDA, ROCm and Metal kernels directly for fast cold starts and a small memory footprint.
A preview-first deployment tool that installs, verifies and revokes custom instruction files for Claude Code across project, local and user scopes without hand-editing CLAUDE.md.
An ingestion pipeline that turns messy unstructured documents into persistent, navigable memory for AI agents — parsing, hierarchy extraction, multimodal structuring and graph construction in one pass.
A fully local memory engine for AI coding assistants, exposed over MCP: verbatim recall with encrypted local storage, benchmarked retrieval quality and injected memory packs that cut search tokens.
Shared, governed memory for fleets of AI agents: agents write plain text, Caura turns it into searchable multi-tenant memory with scoping, trust tiers and cross-agent outcome propagation.
A sandbox-native compute platform for AI agents: stateful Firecracker MicroVMs with snapshots, cloning and live migration, plus a serverless function runtime for long-running orchestration.
A local-first runtime for AI agents providing persistent sessions, sandboxed tool execution, credential vaults, memory, audit trails and a built-in console, all running on your own infrastructure.