Inference, Deployment & Runtime
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
★ 2.2K⑂ 211Python
Apache-2.0Q98
Agents & Multi-Agent
future-agi/future-agi
An open-source, self-hostable platform for evaluating, tracing, simulating, protecting, routing, and optimizing LLM and AI-agent applications.
★ 1.7K⑂ 498Python
Apache-2.0Q98
Inference, Deployment & Runtime
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
★ 1.5K⑂ 279C++
Apache-2.0Q98
Agents & Multi-Agent
0xlazai/alith
A simple, composable, and high-performance AI agent framework designed for Web3 and Crypto, enabling developers to build, deploy, and manage on-chain AI agents with multi-language support and LazAI Gateway integration.
★ 44⑂ 31Rust
Apache-2.0Q88
Inference, Deployment & Runtime
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
★ 314⑂ 35Rust
MITQ92
MCP & Tool Calling
artokun/comfyui-mcp
A local-first, agent-native control plane for ComfyUI that provides an MCP server and sidebar agent to generate images, video, and audio, author and run workflows, and edit live graphs using natural language across any LLM.
★ 599⑂ 93TypeScript
MITQ92
Agents & Multi-Agent
reyamira/models
A TUI and CLI tool for browsing AI models, benchmarks, coding agents, and provider statuses.
★ 492⑂ 19Rust
MITQ92
MCP & Tool Calling
juspay/neurolink
A TypeScript integration platform providing a unified API for 30+ AI providers and 100+ models, enabling provider swapping, multi-modal voice processing, RAG, memory, and MCP-native tool integration.
★ 121⑂ 124TypeScript
MITQ93
MCP & Tool Calling
mtrnix/metronix-memory
Self-hosted memory infrastructure for AI agents featuring MCP-native integration, hybrid RAG, a temporal knowledge graph, and an ontology layer, designed for local-model friendliness and durable, agent-scoped context.
★ 39⑂ 7Python
Apache-2.0Q86
Agents & Multi-Agent
cloudgeni-ai/opengeni
An open, self-hostable agentic runtime for organizations that provides durable, replayable agent sessions, human approvals, governed credentials and memory, and flexible compute targets including managed sandboxes and enrolled user hardware.
★ 56⑂ 5TypeScript
Apache-2.0Q86
MCP & Tool Calling
soapbucket/sbproxy
An open-source, self-hosted Enterprise AI Gateway and LLM proxy written in Rust, providing an OpenAI-compatible API for over 60 providers and supporting local model hosting, MCP tool federation, and robust traffic governance.
★ 49⑂ 1Rust
Apache-2.0Q86
Inference, Deployment & Runtime
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
★ 67⑂ 7Python
Apache-2.0Q89