Awesome Infra for AI › Inference Optimization

nobodywho-ooo/nobodywho

⭐ 1515 Rust repository created 2024-09-03

NobodyWho is an on-device inference engine designed to run Large Language Models (LLMs) and Small Language Models (SLMs) locally and efficiently on a wide range of devices. It eliminates the need for API keys or internet connectivity, supporting offline operation. The engine is compatible with various chat LLMs like Gemma, Qwen, and Mistral, and leverages the GGUF format for model compatibility. A key feature is its fast and type-safe tool calling, which automatically generates structured grammars from function signatures, simplifying integration. NobodyWho also supports multimodal input, allowing for the incorporation of image and audio information into LLM interactions. Users can download models directly from Hugging Face or other URLs. Under the hood, NobodyWho includes conversation-aware preemptive context shifting to maintain full conversation memory without message length limits. It utilizes GPU-accelerated inference via Vulkan or Metal, ensuring high performance across different operating systems. The project provides bindings and documentation for multiple platforms, including Kotlin (Android, Desktop JVM), Swift (iOS, macOS, visionOS, watchOS), Python, Flutter, React Native, and Godot Engine. This broad platform support makes it versatile for integrating local AI capabilities into diverse applications, from mobile games to desktop utilities.

https://github.com/nobodywho-ooo/nobodywho

LLM inferenceon-device AIlocal LLMinference engineGGUFtool callingmultimodal AIGPU accelerationKotlinSwiftPythonFlutterReact NativeGodot

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

alibaba/rtp-llm

RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.