onetoken-oss/K3Flight
Single-file Linux inference server that runs the 2.8T-parameter Kimi K3 on CPU with roughly 55 GB of resident RAM by streaming weights from a local NVMe checkpoint instead of loading it all.
Awesome Infra for AI › Weekly › 2026-08-10
Single-file Linux inference server that runs the 2.8T-parameter Kimi K3 on CPU with roughly 55 GB of resident RAM by streaming weights from a local NVMe checkpoint instead of loading it all.
Rust and CUDA inference engine for giant Mixture-of-Experts models that streams routed experts from NVMe per token, keeping only attention and hot experts in VRAM on consumer GPUs.
Trace-native CI/CD for AI agents: grades every agent trace as it lands, clusters failures into issues, freezes bad runs into hermetic replayable regression cases and blocks the pull request in CI.
Self-hosted multi-provider AI gateway that fronts Gemini, Claude, Codex, Qwen and OpenAI-compatible upstreams behind one endpoint, adding routing groups, failover, request logging, quotas and multi-tenant governance.
LLM inference engine written entirely in Rust and CUDA with no PyTorch or ONNX runtime, serving Qwen3 through trillion-parameter Kimi-K2 over an OpenAI-compatible HTTP API.
Graph-based retrieval engine for LLM memory that scores knowledge units by the strongest evidence path through a four-layer graph instead of ranking chunks by vector similarity alone.
Cloudflare-native agent runtime and control plane that runs Codex, Claude Agent SDK and OpenCode behind API endpoints in isolated sandboxes, with streamed tool activity, durable threads and inspectable runs.