Floe-Labs/floe-guard
In-process spend meter and budget gate for AI agents that hard-stops the next model call, voice turn or paid tool invocation before it crosses a configured cost ceiling.
Awesome Infra for AI › Weekly › 2026-08-24
In-process spend meter and budget gate for AI agents that hard-stops the next model call, voice turn or paid tool invocation before it crosses a configured cost ceiling.
Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.
Reference architecture and TypeScript library that routes AI agent actions through an independent authorization boundary, turning proposals into policy verdicts, human approvals and single-use grants.
Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.
Single-file Linux inference server that runs a 2.8 trillion parameter Kimi K3 checkpoint on CPU by streaming weights from storage, keeping the measured resident working set near 55 GB.
Native MLX port of the DSpark and DFlash speculative decoding drafters, giving lossless multi-fold faster LLM decoding on Apple Silicon with an OpenAI-compatible serving mode.
Trace-native CI/CD for AI agents that grades production traces, clusters failures, freezes bad runs into hermetic replayable regression cases and blocks the pull request that would ship them again.
Self-hosted OpenAI-compatible LLM router that fronts more than a hundred models with bring-your-own-key access, streaming, automatic model selection and an optional hosted fallback.
Self-hosted unified API in front of agent harnesses such as Codex, Claude Code and Hermes, implementing the Unified Harness Protocol with sessions, streaming, files and cancellation.
Local OpenAI- and Anthropic-compatible proxy that exposes a Claude subscription to other AI tools, with session-affinity routing, multi-account pooling and drift detection for long agent runs.
Rust LLM router presenting one OpenAI-compatible endpoint in front of eighty providers and local vLLM or SGLang backends, with cache-affinity routing, RBAC, budgets and no I/O on the request path.
Local-first observability and memory layer for AI coding agents that indexes every session into a typed graph and turns recurring friction into reviewable, tracked experiments.
Self-hosted prompt management service that applies Git-like version control to prompts, exposing them over a REST API with caching, authentication and role-based access.
Self-hosted persistent memory backend for AI agents that stores decisions and context in a semantic store with a typed knowledge graph, served over MCP, REST, CLI and a dashboard.
Agentic retrieval-augmented generation library for local document collections, combining hybrid vector and full-text search, reranking and multimodal retrieval on an embedded LanceDB store.
Structured context layer for AI agents that builds a searchable, continuously refreshed graph of domain entities and relationships from schemas and connectors.