ToolAI.io · GitHub چینل

انفرنس، ڈپلائمنٹ اور رَن ٹائم · GitHub پر اوپن سورس AI پروجیکٹس

LLM، ایجنٹ، MCP، RAG اور AI ڈویلپمنٹ پروجیکٹس کی حقائق پر مبنی فہرست، جس میں ہر اندراج کے لیے لائسنس، سیٹ اپ، ڈاؤن لوڈ، اسکرین شاٹس اور متعلقہ وسائل شامل ہیں۔

عوامی ریپوزٹری کا ڈیٹا

پروجیکٹ انڈیکس

23 پروجیکٹس
Hugging Face Transformers کا اسکرین شاٹ
انفرنس، ڈپلائمنٹ اور رَن ٹائم

Hugging Face Transformers

huggingface/transformers

A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.

★ 163.3K⑂ 34.1KPython
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

vllm

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 89.5K⑂ 21KPython
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

sglang

sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 32.2K⑂ 8.1KPython
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

OpenVINO

openvinotoolkit/openvino

Open-source toolkit for optimizing and deploying AI inference across edge-to-cloud environments.

★ 10.6K⑂ 3.3KC++
Apache-2.0Q98
Xorbits Inference (Xinference) کا اسکرین شاٹ
انفرنس، ڈپلائمنٹ اور رَن ٹائم

Xorbits Inference (Xinference)

xorbitsai/inference

A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.

★ 9.5K⑂ 853Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

Mooncake: KVCache-Centric Infrastructure for Distributed LLM Serving

kvcache-ai/mooncake

Mooncake is a C++ infrastructure project for large-scale LLM inference and training. It separates prefill, decode, and storage resources while providing high-performance transfer, distributed KV-cache storage, and fault-tolerant expert-parallel communication.

★ 6.1K⑂ 1KC++
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

GPUStack

gpustack/gpustack

GPUStack is an open-source GPU cluster manager for serving AI models and provisioning on-demand, SSH-accessible GPU instances. It orchestrates inference engines including vLLM, SGLang, and TensorRT-LLM across on-premises, Kubernetes, and cloud environments.

★ 5.4K⑂ 604Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

llm-d: Distributed LLM Inference on Kubernetes

llm-d/llm-d

llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.

★ 4K⑂ 649Shell
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

FastDeploy

PaddlePaddle/FastDeploy

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

★ 3.7K⑂ 755Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

vLLM Ascend

vllm-project/vllm-ascend

A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.

★ 2.7K⑂ 2.1KC++
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

any-llm: A Unified Python Interface for LLM Providers

mozilla-ai/any-llm

any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.

★ 2.2K⑂ 211Python
Apache-2.0Q98
انفرنس، ڈپلائمنٹ اور رَن ٹائم

llmgateway

theopenco/llmgateway

Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

★ 1.6K⑂ 172TypeScript
NOASSERTIONQ98

صفحہ 1 / 2 · 23 پروجیکٹس

حال ہی میں اپ ڈیٹ کیا گیا

cherry-studioCherryHQ/cherry-studio★ 51.5K siyuansiyuan-note/siyuan★ 46.2K career-opscareer-ops-hq/career-ops★ 70.2K tensorflowtensorflow/tensorflow★ 198.8K streamlitstreamlit/streamlit★ 45.7K pytorchpytorch/pytorch★ 102.8K

سب سے زیادہ اسٹارز والے

ECCaffaan-m/ECC★ 248.8K Hermes Agentnousresearch/hermes-agent★ 227.1K tensorflowtensorflow/tensorflow★ 198.8K AutoGPTSignificant-Gravitas/AutoGPT★ 187.1K ollamaollama/ollama★ 180.2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.9K