ToolAI.io · Chaîne GitHub

Projets d’IA populaires sur GitHub

Suivez les dépôts AI open source actifs dans les domaines des LLM, des agents, de MCP, de RAG et du code, avec leur nombre d’étoiles et des informations vérifiées sur les projets.

Données du dépôt public

Index du projet

24 Projets
Capture d’écran de CSGHub: Open-Source LLM Asset Management Platform
Agents et multi-agents

CSGHub: Open-Source LLM Asset Management Platform

opencsgs/csghub

CSGHub is an open-source, on-premise platform for managing the full lifecycle of Large Language Model assets, including models, datasets, spaces, and code, offering functionality comparable to a private Hugging Face.

★ 4,2K⑂ 525Vue
Apache-2.0Q98
Capture d’écran de Lemonade: Local AI Server for GPU and NPU Inference
MCP et appels d’outils

Lemonade: Local AI Server for GPU and NPU Inference

lemonade-sdk/lemonade

Lemonade is an open-source local AI server that enables users to run optimized Large Language Models (LLMs), speech, and image generation models directly on their own GPUs and NPUs, providing a free and private alternative to cloud APIs.

★ 5,2K⑂ 436C++
Apache-2.0Q98
Inférence, déploiement et exécution

OpenVINO

openvinotoolkit/openvino

Open-source toolkit for optimizing and deploying AI inference across edge-to-cloud environments.

★ 10,6K⑂ 3,3KC++
Apache-2.0Q98
Inférence, déploiement et exécution

Mooncake: KVCache-Centric Infrastructure for Distributed LLM Serving

kvcache-ai/mooncake

Mooncake is a C++ infrastructure project for large-scale LLM inference and training. It separates prefill, decode, and storage resources while providing high-performance transfer, distributed KV-cache storage, and fault-tolerant expert-parallel communication.

★ 6,1K⑂ 1KC++
Apache-2.0Q98
Inférence, déploiement et exécution

llm-d: Distributed LLM Inference on Kubernetes

llm-d/llm-d

llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.

★ 4K⑂ 649Shell
Apache-2.0Q98
Inférence, déploiement et exécution

GPUStack

gpustack/gpustack

GPUStack is an open-source GPU cluster manager for serving AI models and provisioning on-demand, SSH-accessible GPU instances. It orchestrates inference engines including vLLM, SGLang, and TensorRT-LLM across on-premises, Kubernetes, and cloud environments.

★ 5,4K⑂ 604Python
Apache-2.0Q98
Capture d’écran de Qwen3-Coder — Code-focused language models
LLM et modèles de fondation

Qwen3-Coder — Code-focused language models

QwenLM/Qwen3-Coder

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team

★ 16,8K⑂ 1,2KPython
Licence non détectéeQ82

Page 2 / 2 · 24 projets

Récemment mis à jour

codex — Coding agent and developer workflowsopenai/codex★ 106,5K NemoClaw — Agent runtime and deployment toolingNVIDIA/NemoClaw★ 22,2K XERJxerj-org/xerj★ 1,4K LangWatchlangwatch/langwatch★ 3,5K OrchestKityonatangross/orchestkit★ 222 hal0hal0ai/hal0★ 67

Les plus étoilés

Hermes Agentnousresearch/hermes-agent★ 227,1K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174,2K skills — Reusable agent skills and workflowsanthropics/skills★ 169,9K Hugging Face Transformershuggingface/transformers★ 163,3K Firecrawlfirecrawl/firecrawl★ 161,1K LangChainlangchain-ai/langchain★ 143,6K