ToolAI.io · GitHub चैनल

इन्फरेंस, डिप्लॉयमेंट और रनटाइम · GitHub पर ओपन-सोर्स AI प्रोजेक्ट्स

LLM, एजेंट, MCP, RAG और AI डेवलपमेंट प्रोजेक्ट्स की तथ्यों पर आधारित इंडेक्स, जिसमें हर प्रविष्टि के लिए लाइसेंस, सेटअप, डाउनलोड, स्क्रीनशॉट और संबंधित संसाधन शामिल हैं।

सार्वजनिक रिपॉज़िटरी का डेटा

प्रोजेक्ट इंडेक्स

23 प्रोजेक्ट्स
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

xLLM: High-Performance Inference Engine for Diverse AI Accelerators

xllm-ai/xllm

xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.

★ 1.5K⑂ 279C++
Apache-2.0Q98
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

TensorRT

NVIDIA/TensorRT

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

★ 13.2K⑂ 2.4KC++
Apache-2.0Q94
memra का स्क्रीनशॉट
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

memra

avifenesh/memra

A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.

★ 314⑂ 35Rust
MITQ92
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

Paddle-Lite

PaddlePaddle/Paddle-Lite

PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)

★ 7.3K⑂ 1.6KC++
Apache-2.0Q91
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

ollama

ollama/ollama

A tool for running and managing large language models locally.

★ 180.2K⑂ 17.7KGo
MITQ90
hal0 का स्क्रीनशॉट
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

hal0

hal0ai/hal0

An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.

★ 68⑂ 7Python
Apache-2.0Q89
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

Model-Optimizer

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

★ 3.3K⑂ 513Python
Apache-2.0Q86
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

LocalAI

mudler/LocalAI

A local inference engine for self-hosting models and AI services.

★ 48.9K⑂ 4.4KGo
MITQ85
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

ray

ray-project/ray

A distributed computing runtime for machine learning training, tuning and model serving.

★ 43.7K⑂ 8KPython
Apache-2.0Q85
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

checkpoint-engine

MoonshotAI/checkpoint-engine

Checkpoint-engine is a simple middleware to update model weights in LLM inference engines

★ 982⑂ 100Python
MITQ83
इन्फरेंस, डिप्लॉयमेंट और रनटाइम

jan

janhq/jan

A local AI chat application for running models on a personal computer.

★ 44.3K⑂ 3KTypeScript
लाइसेंस का पता नहीं चलाQ80

पृष्ठ 2 / 2 · 23 प्रोजेक्ट

हाल ही में अपडेट किए गए

cherry-studioCherryHQ/cherry-studio★ 51.5K siyuansiyuan-note/siyuan★ 46.2K career-opscareer-ops-hq/career-ops★ 70.2K tensorflowtensorflow/tensorflow★ 198.8K streamlitstreamlit/streamlit★ 45.7K pytorchpytorch/pytorch★ 102.8K

सर्वाधिक स्टार वाले

ECCaffaan-m/ECC★ 248.8K Hermes Agentnousresearch/hermes-agent★ 227.1K tensorflowtensorflow/tensorflow★ 198.8K AutoGPTSignificant-Gravitas/AutoGPT★ 187.1K ollamaollama/ollama★ 180.2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.9K