Awesome Infra for AI › Inference Optimization

PaddlePaddle/Paddle.js

⭐ 1106 JavaScript repository created 2020-03-26

Paddle.js is a specialized web project built for Baidu PaddlePaddle, designed to facilitate the execution of deep learning models directly within web browsers and mini-programs. Its core functionality revolves around providing an inference engine that leverages web technologies like WebGL, WebGPU, and WebAssembly, enabling high-performance model predictions on client-side devices. The project supports loading pre-trained models and includes tools for converting models from PaddlePaddle's ecosystem into a browser-friendly format. This capability allows developers to deploy AI/ML models without requiring server-side inference, enhancing interactivity and reducing latency for applications such as computer vision tasks (e.g., human segmentation, OCR, gesture recognition, face detection). Paddle.js comprises several key modules, including `paddlejs-core` for managing the inference process and various backend implementations (`webgl`, `wasm`, `webgpu`, `cpu`) to optimize performance across different hardware and browser capabilities. It also offers a model transformation tool, `paddlejsconverter`, to prepare models for browser deployment. The ecosystem extends to specific pre-built models and examples, making it easier for developers to integrate AI functionalities into web-based applications.

https://github.com/PaddlePaddle/Paddle.js

deep-learninginference-enginewebwebassemblywebglwebgpumodel servingbrowser inferencepaddlepaddle

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.