ToolAI.io · GitHub Channel

Inference, Deployment & Runtime · Open-source AI Projects on GitHub

A fact-based index of LLM, agent, MCP, RAG and AI development projects, with license, setup, download, screenshots and related resources for each entry.

Public repository data

Project index

23 projects
Inference, Deployment & Runtime

xLLM: High-Performance Inference Engine for Diverse AI Accelerators

xllm-ai/xllm

xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.

★ 1.5K⑂ 279C++
Apache-2.0Q98
Inference, Deployment & Runtime

TensorRT

NVIDIA/TensorRT

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

★ 13.2K⑂ 2.4KC++
Apache-2.0Q94
Screenshot of memra
Inference, Deployment & Runtime

memra

avifenesh/memra

A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.

★ 314⑂ 35Rust
MITQ92
Inference, Deployment & Runtime

Paddle-Lite

PaddlePaddle/Paddle-Lite

PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)

★ 7.3K⑂ 1.6KC++
Apache-2.0Q91
Inference, Deployment & Runtime

ollama

ollama/ollama

A tool for running and managing large language models locally.

★ 180.2K⑂ 17.7KGo
MITQ90
Screenshot of hal0
Inference, Deployment & Runtime

hal0

hal0ai/hal0

An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.

★ 68⑂ 7Python
Apache-2.0Q89
Inference, Deployment & Runtime

Model-Optimizer

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

★ 3.3K⑂ 513Python
Apache-2.0Q86
Inference, Deployment & Runtime

LocalAI

mudler/LocalAI

A local inference engine for self-hosting models and AI services.

★ 48.9K⑂ 4.4KGo
MITQ85
Inference, Deployment & Runtime

ray

ray-project/ray

A distributed computing runtime for machine learning training, tuning and model serving.

★ 43.7K⑂ 8KPython
Apache-2.0Q85
Inference, Deployment & Runtime

checkpoint-engine

MoonshotAI/checkpoint-engine

Checkpoint-engine is a simple middleware to update model weights in LLM inference engines

★ 982⑂ 100Python
MITQ83
Inference, Deployment & Runtime

jan

janhq/jan

A local AI chat application for running models on a personal computer.

★ 44.3K⑂ 3KTypeScript
License not detectedQ80

Page 2 / 2 · 23 projects

Recently updated

cherry-studioCherryHQ/cherry-studio★ 51.5K siyuansiyuan-note/siyuan★ 46.2K career-opscareer-ops-hq/career-ops★ 70.2K tensorflowtensorflow/tensorflow★ 198.8K streamlitstreamlit/streamlit★ 45.7K pytorchpytorch/pytorch★ 102.8K

Most starred

ECCaffaan-m/ECC★ 248.8K Hermes Agentnousresearch/hermes-agent★ 227.1K tensorflowtensorflow/tensorflow★ 198.8K AutoGPTSignificant-Gravitas/AutoGPT★ 187.1K ollamaollama/ollama★ 180.2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.9K