ToolAI.io · Chaîne GitHub

Nouveaux projets AI sur GitHub

Découvrez les nouveaux projets AI open source ajoutés sur GitHub, avec les informations du dépôt, les instructions de configuration, les licences et les ressources associées.

Données du dépôt public

Index du projet

24 Projets
Inférence, déploiement et exécution

xLLM: High-Performance Inference Engine for Diverse AI Accelerators

xllm-ai/xllm

xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.

★ 1,5K⑂ 279C++
Apache-2.0Q98
Capture d’écran de memra
Inférence, déploiement et exécution

memra

avifenesh/memra

A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.

★ 314⑂ 35Rust
MITQ92
Capture d’écran de hal0
Inférence, déploiement et exécution

hal0

hal0ai/hal0

An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.

★ 67⑂ 7Python
Apache-2.0Q89
Capture d’écran de Qwen3-Coder — Code-focused language models
LLM et modèles de fondation

Qwen3-Coder — Code-focused language models

QwenLM/Qwen3-Coder

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team

★ 16,8K⑂ 1,2KPython
Licence non détectéeQ82
Capture d’écran de TensorRT-LLM — Optimized LLM inference
Inférence, déploiement et exécution

TensorRT-LLM — Optimized LLM inference

NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way

★ 14,4K⑂ 2,7KC++
NOASSERTIONQ82
Capture d’écran de tiktoken — Fast tokenization
Inférence, déploiement et exécution

tiktoken — Fast tokenization

openai/tiktoken

tiktoken is a fast BPE tokeniser for use with OpenAI's models

★ 19K⑂ 1,6KRust
MITQ86

Page 2 / 2 · 24 projets

Récemment mis à jour

qwen-code — Command-line coding agentQwenLM/qwen-code★ 27,1K SBproxysoapbucket/sbproxy★ 49 XERJxerj-org/xerj★ 1,4K PwrAgentpwrdrvr/pwragent★ 29 OpenGenicloudgeni-ai/opengeni★ 56 NemoClaw — Agent runtime and deployment toolingNVIDIA/NemoClaw★ 22,2K

Les plus étoilés

Hermes Agentnousresearch/hermes-agent★ 227,1K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174,3K skills — Reusable agent skills and workflowsanthropics/skills★ 170,1K Hugging Face Transformershuggingface/transformers★ 163,3K Firecrawlfirecrawl/firecrawl★ 161,1K LangChainlangchain-ai/langchain★ 143,6K