vLLM Ascend
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
ToolAI.io · Canal do GitHub
Navegue pelo diretório completo da ToolAI de projetos de AI de código aberto no GitHub, organizado por tópico, linguagem e licença.
Dados do repositório público
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
bionic-gpt/bionic-gpt
Bionic is a Rust-based, on-premise alternative to ChatGPT designed for private generative AI deployments. It provides a familiar chat interface, local or remote model access, team controls, retrieval-augmented assistants, data integrations, observability, and Kubernetes-oriented scaling.
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
future-agi/future-agi
An open-source, self-hostable platform for evaluating, tracing, simulating, protecting, routing, and optimizing LLM and AI-agent applications.
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
askimo-ai/askimo
A native desktop AI client for chat, local RAG, multi-step AI workflows (Plans), and agent skills, supporting multiple cloud and local LLM providers while keeping user files strictly on the machine.
juspay/neurolink
A TypeScript integration platform providing a unified API for 30+ AI providers and 100+ models, enabling provider swapping, multi-modal voice processing, RAG, memory, and MCP-native tool integration.
artokun/comfyui-mcp
A local-first, agent-native control plane for ComfyUI that provides an MCP server and sidebar agent to generate images, video, and audio, author and run workflows, and edit live graphs using natural language across any LLM.
reyamira/models
A TUI and CLI tool for browsing AI models, benchmarks, coding agents, and provider statuses.
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
sno-ai/llmix
A production LLM call layer for AI agents and tools that wraps existing provider SDKs with config-driven model presets, caching, resilience patterns, and key rotation across Python, TypeScript, and Rust.
taichuy/1flowbase
An open-source AI gateway that allows local agent clients to publish fusion-style multi-model workflows as OpenAI and Claude-compatible virtual models, providing full observability into traces, tokens, latency, and costs.
Página 5 / 7 · 75 projetos