Awesome Infra for AI › Inference Optimization

aiptimizer/TurboOCR

⭐ 1104 C++ repository created 2026-03-20

TurboOCR is a GPU-accelerated server application specializing in Optical Character Recognition (OCR), offering significantly faster performance compared to Python-based solutions. It achieves high throughput (up to 270 img/s on FUNSD dataset) and low latency (11 ms p50) by utilizing C++, CUDA, and NVIDIA TensorRT for FP16 inference, based on the PaddleOCR PP-OCRv5 model. The server supports OCR for both printed and handwritten text, offers various PDF processing modes (pure OCR, native text layer, auto-dispatch, detection-verified hybrid), and includes advanced features like layout detection (PP-DocLayoutV3 with 25 region classes) and intelligent reading order reconstruction. It provides both HTTP and gRPC APIs from a single binary, sharing a common GPU pipeline pool, making it suitable for high-demand AI inference serving. Deployment is simplified with a one-line Docker command, which automatically builds TensorRT engines. For operational visibility, TurboOCR integrates Prometheus metrics, exposing request counters, latency histograms, and VRAM usage. It supports multiple languages including Latin script languages, Chinese, Greek, Russian, Arabic, Korean, and Thai, and is designed for Linux systems with NVIDIA GPUs (Turing or newer). Planned features include structured extraction and table parsing, indicating its ongoing development towards comprehensive document AI.

https://github.com/aiptimizer/TurboOCR

document-aidocument-parsinggpu-ocrinference-serverocrpaddleocrpdf-extractiontensorrttext-detectiontext-recognitionfastapigrpcprometheusdockerproduction-readyai-serving

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.