Inférence, déploiement et exécution
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
★ 1,5K⑂ 279C++
Apache-2.0Q98
Inférence, déploiement et exécution
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
★ 314⑂ 35Rust
MITQ92
Inférence, déploiement et exécution
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
★ 67⑂ 7Python
Apache-2.0Q89
LLM et modèles de fondation
À la une
deepseek-ai/DeepSeek-V3
An open-source project focused on large language model research and inference.
★ 104,3K⑂ 16,7KPython
MITQ86
LLM et modèles de fondation
À la une
deepseek-ai/DeepSeek-R1
An open-source project focused on reasoning model research and evaluation.
★ 92K⑂ 11,7KPython
MITQ86
LLM et modèles de fondation
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
★ 27,5K⑂ 2KPython
Licence non détectéeQ82
LLM et modèles de fondation
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
★ 16,8K⑂ 1,2KPython
Licence non détectéeQ82
LLM et modèles de fondation
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
★ 24,2K⑂ 2,9KPython
MITQ86
Inférence, déploiement et exécution
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way
★ 14,4K⑂ 2,7KC++
NOASSERTIONQ82
Inférence, déploiement et exécution
openai/tiktoken
tiktoken is a fast BPE tokeniser for use with OpenAI's models
★ 19K⑂ 1,6KRust
MITQ86
LLM et modèles de fondation
openai/gpt-oss
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20,3K⑂ 2,1KPython
Apache-2.0Q86
Agents et multi-agents
NVIDIA/NemoClaw
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
★ 22,2K⑂ 3KPython
Apache-2.0Q86