Awesome Infra for AI › Inference Optimization

Tencent/Forward

⭐ 557 C++ repository created 2021-03-11

Forward is a high-performance deep learning inference acceleration framework designed by Tencent, specifically for NVIDIA GPUs. It simplifies the process of deploying trained models from popular frameworks like TensorFlow, PyTorch, Keras, and ONNX by converting them into optimized TensorRT inference engines. This eliminates the need for complex intermediate model conversion steps or network construction, making it more user-friendly and extensible than directly using TensorRT. The framework focuses on achieving superior inference performance through network-level optimization and supports various advanced models in computer vision (CV), natural language processing (NLP), and recommendation systems, including BERT, FaceSwap, and StyleTransfer. Key features of Forward include high model performance optimization, broad model support, multiple inference modes (FLOAT, HALF, INT8), a simple C++ and Python API for direct import of trained models, and extensibility for custom network layers. It requires NVIDIA CUDA, CuDNN, and TensorRT, along with specific versions of PyTorch, TensorFlow, and CMake. The project provides clear documentation and examples for building and using the framework in both C++ and Python, making it accessible for developers looking to accelerate deep learning inference in production environments.

https://github.com/Tencent/Forward

deep-learning inferenceGPU accelerationTensorRTmodel servinginference optimizationTensorFlow inferencePyTorch inferenceKeras inferenceONNX inference

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.