ToolAI.io · Chaîne GitHub

Projets d’IA populaires sur GitHub

Suivez les dépôts AI open source actifs dans les domaines des LLM, des agents, de MCP, de RAG et du code, avec leur nombre d’étoiles et des informations vérifiées sur les projets.

Données du dépôt public

Index du projet

24 Projets
Capture d’écran de hal0
Inférence, déploiement et exécution

hal0

hal0ai/hal0

An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.

★ 67⑂ 7Python
Apache-2.0Q89
Capture d’écran de memra
Inférence, déploiement et exécution

memra

avifenesh/memra

A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.

★ 312⑂ 35Rust
MITQ92
Capture d’écran de TensorRT-LLM — Optimized LLM inference
Inférence, déploiement et exécution

TensorRT-LLM — Optimized LLM inference

NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way

★ 14,4K⑂ 2,7KC++
NOASSERTIONQ82
Agents et multi-agents

dstack

dstackai/dstack

A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.

★ 2,2K⑂ 248Python
MPL-2.0Q98
Inférence, déploiement et exécution

vLLM Ascend

vllm-project/vllm-ascend

A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.

★ 2,7K⑂ 2,1KC++
Apache-2.0Q98
Inférence, déploiement et exécution

any-llm: A Unified Python Interface for LLM Providers

mozilla-ai/any-llm

any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.

★ 2,2K⑂ 211Python
Apache-2.0Q98
Inférence, déploiement et exécution

xLLM: High-Performance Inference Engine for Diverse AI Accelerators

xllm-ai/xllm

xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.

★ 1,5K⑂ 279C++
Apache-2.0Q98
Capture d’écran de tiktoken — Fast tokenization
Inférence, déploiement et exécution

tiktoken — Fast tokenization

openai/tiktoken

tiktoken is a fast BPE tokeniser for use with OpenAI's models

★ 19K⑂ 1,6KRust
MITQ86
Capture d’écran de Xorbits Inference (Xinference)
Inférence, déploiement et exécution

Xorbits Inference (Xinference)

xorbitsai/inference

A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.

★ 9,5K⑂ 853Python
Apache-2.0Q98
Capture d’écran de Plano: AI-Native Proxy Server and Data Plane for Agentic Apps
Agents et multi-agents

Plano: AI-Native Proxy Server and Data Plane for Agentic Apps

katanemo/plano

Plano is an AI-native proxy server and data plane built in Rust that centralizes LLM routing, agent orchestration, observability, and guardrails, allowing developers to focus on core agent logic rather than infrastructure plumbing.

★ 7K⑂ 484Rust
Apache-2.0Q98
Capture d’écran de Hugging Face Transformers
Inférence, déploiement et exécution

Hugging Face Transformers

huggingface/transformers

A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.

★ 163,3K⑂ 34,1KPython
Apache-2.0Q98

Page 1 / 2 · 24 projets

Récemment mis à jour

codex — Coding agent and developer workflowsopenai/codex★ 106,5K NemoClaw — Agent runtime and deployment toolingNVIDIA/NemoClaw★ 22,2K XERJxerj-org/xerj★ 1,4K LangWatchlangwatch/langwatch★ 3,5K OrchestKityonatangross/orchestkit★ 222 hal0hal0ai/hal0★ 67

Les plus étoilés

Hermes Agentnousresearch/hermes-agent★ 227,1K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174,2K skills — Reusable agent skills and workflowsanthropics/skills★ 169,9K Hugging Face Transformershuggingface/transformers★ 163,3K Firecrawlfirecrawl/firecrawl★ 161,1K LangChainlangchain-ai/langchain★ 143,6K