ToolAI.io · GitHub چینل

GitHub پر نئے AI پروجیکٹس

GitHub پر حال ہی میں شامل کیے گئے اوپن سورس AI پروجیکٹس دریافت کریں، جن کے ساتھ ریپوزٹری کی معلومات، سیٹ اپ نوٹس، لائسنس اور متعلقہ وسائل بھی موجود ہیں۔

عوامی ریپوزٹری کا ڈیٹا

پروجیکٹ انڈیکس

41 پروجیکٹس
انفرنس، ڈپلائمنٹ اور رَن ٹائم

GPUStack

gpustack/gpustack

GPUStack is an open-source GPU cluster manager for serving AI models and provisioning on-demand, SSH-accessible GPU instances. It orchestrates inference engines including vLLM, SGLang, and TensorRT-LLM across on-premises, Kubernetes, and cloud environments.

★ 5.4K⑂ 604Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

llm-d: Distributed LLM Inference on Kubernetes

llm-d/llm-d

llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.

★ 4K⑂ 649Shell
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

vLLM Ascend

vllm-project/vllm-ascend

A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.

★ 2.7K⑂ 2.1KC++
Apache-2.0Q98
ایجنٹس اور ملٹی ایجنٹ

dstack

dstackai/dstack

A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.

★ 2.2K⑂ 249Python
MPL-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

any-llm: A Unified Python Interface for LLM Providers

mozilla-ai/any-llm

any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.

★ 2.2K⑂ 211Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

xLLM: High-Performance Inference Engine for Diverse AI Accelerators

xllm-ai/xllm

xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.

★ 1.5K⑂ 279C++
Apache-2.0Q98
memra کا اسکرین شاٹ
انفرنس، ڈپلائمنٹ اور رَن ٹائم

memra

avifenesh/memra

A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.

★ 314⑂ 35Rust
MITQ92
hal0 کا اسکرین شاٹ
انفرنس، ڈپلائمنٹ اور رَن ٹائم

hal0

hal0ai/hal0

An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.

★ 68⑂ 7Python
Apache-2.0Q89
Qwen3-Coder — Code-focused language models کا اسکرین شاٹ
LLM اور فاؤنڈیشن ماڈلز

Qwen3-Coder — Code-focused language models

QwenLM/Qwen3-Coder

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team

★ 16.8K⑂ 1.2KPython
لائسنس کا پتا نہیں چلاQ82

صفحہ 3 / 4 · 41 پروجیکٹس

حال ہی میں اپ ڈیٹ کیا گیا

cherry-studioCherryHQ/cherry-studio★ 51.5K siyuansiyuan-note/siyuan★ 46.2K career-opscareer-ops-hq/career-ops★ 70.2K tensorflowtensorflow/tensorflow★ 198.8K streamlitstreamlit/streamlit★ 45.7K pytorchpytorch/pytorch★ 102.8K

سب سے زیادہ اسٹارز والے

ECCaffaan-m/ECC★ 248.8K Hermes Agentnousresearch/hermes-agent★ 227.1K tensorflowtensorflow/tensorflow★ 198.8K AutoGPTSignificant-Gravitas/AutoGPT★ 187.1K ollamaollama/ollama★ 180.2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.9K