Awesome Infra for AI › Weekly › 2026-09-28

2026-09-28

9 projects added

AI Safety & Guardrails

hoophq/hoop

Open-source sidecar that sits in front of databases and other resources to mask sensitive data and block destructive queries before AI agents can run them.

Inference Optimization

carloslfu/slotstream

Native Swift inference engine that runs 100GB+ open LLMs on Macs with only 16-64GB of memory by streaming model weights from SSD as needed.

LLM Evaluation & Testing

hermes-labs-ai/lintlang

Local, deterministic static linter that scans agent instructions, MCP tool schemas and prompt files for ambiguity, conflicts and missing bounds before an agent runs.

LLM Gateways & Proxies

apache/casbin-gateway

Local gateway and dashboard that routes every AI coding agent on a machine through one policy layer, tracking spend, permissions and provider authenticity across 44 model vendors.

wink-run/tokenbank

Local gateway that puts Claude Code, Cursor, Codex and other agent CLIs behind one endpoint, tracing usage and routing requests across local models, quotas and paid APIs.

LLM Observability & Tracing

LockedinLabs-AI/agent-console

Local-first console that aggregates token usage, cache activity, model choice and estimated cost across Claude Code and Codex sessions on every connected machine.

splunk/token-meter

Local-first dashboard that turns Claude, Codex, Cursor and other coding-agent session logs into one view of token usage, estimated cost and context pressure.

Model Serving Frameworks

raullenchai/Rapid-MLX

OpenAI- and Anthropic-compatible LLM inference server and native Mac app for Apple Silicon, built on MLX and optimized for reliable tool calling by coding agents.

Vector Databases & Retrieval Infrastructure

LibreChat-AI/rag-api

FastAPI service that indexes documents into pgvector embeddings by file ID and serves scoped retrieval queries for RAG applications, built primarily for LibreChat.

Older issue