NemoClaw — Agent runtime and deployment tooling
NVIDIA/NemoClaw
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
ToolAI.io · Canale GitHub
Monitora repository AI open source attivi nei temi LLM, agenti, MCP, RAG e programmazione, con il numero di stelle e dati verificati sui progetti.
Dati del repository pubblico
NVIDIA/NemoClaw
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
openai/tiktoken
tiktoken is a fast BPE tokeniser for use with OpenAI's models
xorbitsai/inference
A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.
katanemo/plano
Plano is an AI-native proxy server and data plane built in Rust that centralizes LLM routing, agent orchestration, observability, and guardrails, allowing developers to focus on core agent logic rather than infrastructure plumbing.
huggingface/transformers
A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.
Pagina 1 / 2 · 24 progetti