dstack
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
ToolAI.io · Kênh GitHub
Duyệt toàn bộ thư mục ToolAI gồm các dự án AI mã nguồn mở trên GitHub, được sắp xếp theo chủ đề, ngôn ngữ và giấy phép.
Dữ liệu kho lưu trữ công khai
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
NVIDIA/NemoClaw
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
openai/gpt-oss
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
openai/tiktoken
tiktoken is a fast BPE tokeniser for use with OpenAI's models
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way
Trang 2 / 2 · 24 dự án