dstack
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
ToolAI.io · GitHub-Kanal
Durchsuchen Sie das vollständige ToolAI-Verzeichnis mit Open-Source-AI-Projekten auf GitHub, geordnet nach Thema, Sprache und Lizenz.
Daten des öffentlichen Repositorys
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
NVIDIA/NemoClaw
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
openai/gpt-oss
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
openai/tiktoken
tiktoken is a fast BPE tokeniser for use with OpenAI's models
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way
Seite 2 / 2 · 24 Projekte