ToolAI.io · Canal do GitHub

Inferência, implantação e execução · Projetos de AI de código aberto no GitHub

Um índice baseado em fatos de projetos de desenvolvimento de LLM, agentes, MCP, RAG e AI, com licença, configuração, download, capturas de tela e recursos relacionados para cada item.

Dados do repositório público

Índice do projeto

24 Projetos
Captura de tela de Hugging Face Transformers
Inferência, implantação e execução

Hugging Face Transformers

huggingface/transformers

A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.

★ 163,3K⑂ 34,1KPython
Apache-2.0Q98
Inferência, implantação e execução

vllm

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 89,5K⑂ 21KPython
Apache-2.0Q98
Inferência, implantação e execução

sglang

sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 32,2K⑂ 8,1KPython
Apache-2.0Q98
Inferência, implantação e execução

OpenVINO

openvinotoolkit/openvino

Open-source toolkit for optimizing and deploying AI inference across edge-to-cloud environments.

★ 10,6K⑂ 3,3KC++
Apache-2.0Q98
Captura de tela de Xorbits Inference (Xinference)
Inferência, implantação e execução

Xorbits Inference (Xinference)

xorbitsai/inference

A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.

★ 9,5K⑂ 853Python
Apache-2.0Q98
Inferência, implantação e execução

Mooncake: KVCache-Centric Infrastructure for Distributed LLM Serving

kvcache-ai/mooncake

Mooncake is a C++ infrastructure project for large-scale LLM inference and training. It separates prefill, decode, and storage resources while providing high-performance transfer, distributed KV-cache storage, and fault-tolerant expert-parallel communication.

★ 6,1K⑂ 1KC++
Apache-2.0Q98
Inferência, implantação e execução

GPUStack

gpustack/gpustack

GPUStack is an open-source GPU cluster manager for serving AI models and provisioning on-demand, SSH-accessible GPU instances. It orchestrates inference engines including vLLM, SGLang, and TensorRT-LLM across on-premises, Kubernetes, and cloud environments.

★ 5,4K⑂ 604Python
Apache-2.0Q98
Inferência, implantação e execução

llm-d: Distributed LLM Inference on Kubernetes

llm-d/llm-d

llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.

★ 4K⑂ 649Shell
Apache-2.0Q98
Inferência, implantação e execução

FastDeploy

PaddlePaddle/FastDeploy

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

★ 3,7K⑂ 755Python
Apache-2.0Q98
Inferência, implantação e execução

vLLM Ascend

vllm-project/vllm-ascend

A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.

★ 2,7K⑂ 2,1KC++
Apache-2.0Q98
Inferência, implantação e execução

any-llm: A Unified Python Interface for LLM Providers

mozilla-ai/any-llm

any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.

★ 2,2K⑂ 211Python
Apache-2.0Q98
Inferência, implantação e execução

llmgateway

theopenco/llmgateway

Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

★ 1,6K⑂ 172TypeScript
NOASSERTIONQ98

Página 1 / 2 · 24 projetos

Atualizado recentemente

TypeSafe Python SDKtypesafe-ai/typesafe-sdk-python★ 218 TypeSafe JavaScript SDKtypesafe-ai/typesafe-sdk-js★ 231 TypeSafe Agent Skillstypesafe-ai/skills★ 2K cherry-studioCherryHQ/cherry-studio★ 51,5K onyxonyx-dot-app/onyx★ 32K siyuansiyuan-note/siyuan★ 46,2K

Mais estrelados

ECCaffaan-m/ECC★ 248,8K Hermes Agentnousresearch/hermes-agent★ 227,1K tensorflowtensorflow/tensorflow★ 198,8K AutoGPTSignificant-Gravitas/AutoGPT★ 187,1K ollamaollama/ollama★ 180,2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174,9K