ToolAI.io · GitHub چینل

GitHub پر AI پروجیکٹ ڈائریکٹری

GitHub پر موضوع، زبان اور لائسنس کے لحاظ سے منظم اوپن سورس AI پروجیکٹس کی مکمل ToolAI ڈائریکٹری براؤز کریں۔

عوامی ریپوزٹری کا ڈیٹا

پروجیکٹ انڈیکس

41 پروجیکٹس
انفرنس، ڈپلائمنٹ اور رَن ٹائم

llm-d: Distributed LLM Inference on Kubernetes

llm-d/llm-d

llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.

★ 4K⑂ 649Shell
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

FastDeploy

PaddlePaddle/FastDeploy

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

★ 3.7K⑂ 755Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

vLLM Ascend

vllm-project/vllm-ascend

A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.

★ 2.7K⑂ 2.1KC++
Apache-2.0Q98
ایجنٹس اور ملٹی ایجنٹ

dstack

dstackai/dstack

A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.

★ 2.2K⑂ 249Python
MPL-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

any-llm: A Unified Python Interface for LLM Providers

mozilla-ai/any-llm

any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.

★ 2.2K⑂ 211Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

llmgateway

theopenco/llmgateway

Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

★ 1.6K⑂ 172TypeScript
NOASSERTIONQ98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

xLLM: High-Performance Inference Engine for Diverse AI Accelerators

xllm-ai/xllm

xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.

★ 1.5K⑂ 279C++
Apache-2.0Q98
فائن ٹیوننگ، تربیت اور ڈیٹا

DeepSpeed

deepspeedai/deepspeed

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

★ 43K⑂ 4.9KPython
Apache-2.0Q94
انفرنس، ڈپلائمنٹ اور رَن ٹائم

TensorRT

NVIDIA/TensorRT

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

★ 13.2K⑂ 2.4KC++
Apache-2.0Q94
memra کا اسکرین شاٹ
انفرنس، ڈپلائمنٹ اور رَن ٹائم

memra

avifenesh/memra

A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.

★ 314⑂ 35Rust
MITQ92
انفرنس، ڈپلائمنٹ اور رَن ٹائم

Paddle-Lite

PaddlePaddle/Paddle-Lite

PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)

★ 7.3K⑂ 1.6KC++
Apache-2.0Q91
انفرنس، ڈپلائمنٹ اور رَن ٹائم

ollama

ollama/ollama

A tool for running and managing large language models locally.

★ 180.2K⑂ 17.7KGo
MITQ90

صفحہ 2 / 4 · 41 پروجیکٹس

حال ہی میں اپ ڈیٹ کیا گیا

cherry-studioCherryHQ/cherry-studio★ 51.5K siyuansiyuan-note/siyuan★ 46.2K career-opscareer-ops-hq/career-ops★ 70.2K tensorflowtensorflow/tensorflow★ 198.8K streamlitstreamlit/streamlit★ 45.7K pytorchpytorch/pytorch★ 102.8K

سب سے زیادہ اسٹارز والے

ECCaffaan-m/ECC★ 248.8K Hermes Agentnousresearch/hermes-agent★ 227.1K tensorflowtensorflow/tensorflow★ 198.8K AutoGPTSignificant-Gravitas/AutoGPT★ 187.1K ollamaollama/ollama★ 180.2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.9K