Awesome Infra for AI › Inference Optimization

Tencent/FeatherCNN

⭐ 1228 C++ repository created 2018-04-27

FeatherCNN, developed by Tencent AI Platform Department, is designed as a high-performance, lightweight inference engine primarily for convolutional neural networks (CNNs). Its core purpose is to enable efficient execution of trained AI models on resource-constrained platforms, specifically targeting ARM CPUs found in mobile phones (iOS/Android), embedded devices (Linux), and ARM-based servers. The project originated from a game AI initiative for 'King of Glory,' highlighting its practical application in deploying neural models on mobile devices. FeatherCNN distinguishes itself with several key features: it delivers state-of-the-art inference computing performance across a wide range of devices, ensures easy deployment by consolidating everything into a single code base without third-party dependencies, and maintains a featherweight compiled library size (hundreds of KBs). It accepts Caffe models, converting them into a proprietary '.feathermodel' format for optimized inference. The library provides basic C++ runtime interfaces for initializing networks and performing forward computations, with emphasis on raw pointer usage for data handling. While it shares some concepts with general inference frameworks, its explicit focus on lightweight and high-performance inference for CNNs on specific architectures (ARM) makes it a specialized tool for model serving and optimization rather than a general-purpose ML library. The project's emphasis is purely on the operational aspect of already-trained models.

https://github.com/Tencent/FeatherCNN

inference engineCNNmobile AIARM optimizationdeep learningmodel servingembedded AIhigh performancelightweightCaffe

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.