Awesome Infra for AI › Weekly › 2026-08-10

2026-08-10

7 projects added

Inference Optimization

onetoken-oss/K3Flight

Single-file Linux inference server that runs the 2.8T-parameter Kimi K3 on CPU with roughly 55 GB of resident RAM by streaming weights from a local NVMe checkpoint instead of loading it all.

giannisanni/pulsar

Rust and CUDA inference engine for giant Mixture-of-Experts models that streams routed experts from NVMe per token, keeping only attention and hot experts in VRAM on consumer GPUs.

LLM Evaluation & Testing

Jwuthri/Tracely

Trace-native CI/CD for AI agents: grades every agent trace as it lands, clusters failures into issues, freezes bad runs into hermetic replayable regression cases and blocks the pull request in CI.

LLM Gateways & Proxies

kittors/CliRelay

Self-hosted multi-provider AI gateway that fronts Gemini, Claude, Codex, Qwen and OpenAI-compatible upstreams behind one endpoint, adding routing groups, failover, request logging, quotas and multi-tenant governance.

Model Serving Frameworks

pegainfer-project/pegainfer

LLM inference engine written entirely in Rust and CUDA with no PyTorch or ONNX runtime, serving Qwen3 through trillion-parameter Kimi-K2 over an OpenAI-compatible HTTP API.

Vector Databases & Retrieval Infrastructure

FlowElement-xinliuyuansu/m_flow

Graph-based retrieval engine for LLM memory that scores knowledge units by the strongest evidence path through a four-layer graph instead of ranking chunks by vector similarity alone.

Workflow Orchestration for AI

langgenius/mosoo

Cloudflare-native agent runtime and control plane that runs Codex, Claude Agent SDK and OpenCode behind API endpoints in isolated sandboxes, with streamed tool activity, durable threads and inspectable runs.

Newer issue Older issue