Inferentie, implementatie en runtime
llm-d/llm-d
llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.
★ 4K⑂ 649Shell
Apache-2.0Q98
Agents en multi-agents
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
★ 3,5K⑂ 345TypeScript
Apache-2.0Q98
Inferentie, implementatie en runtime
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
★ 2,7K⑂ 2,1KC++
Apache-2.0Q98
LLM en foundationmodellen
bionic-gpt/bionic-gpt
Bionic is a Rust-based, on-premise alternative to ChatGPT designed for private generative AI deployments. It provides a familiar chat interface, local or remote model access, team controls, retrieval-augmented assistants, data integrations, observability, and Kubernetes-oriented scaling.
★ 2,4K⑂ 239Rust
NOASSERTIONQ98
Agents en multi-agents
embabel/embabel-agent
Embabel is an open-source JVM framework for building strongly typed agentic applications that combine LLM interactions, regular code, tools, and domain models. It supports dynamic planning and can be used from Kotlin or Java.
★ 4,3K⑂ 417Kotlin
Apache-2.0Q98
Agents en multi-agents
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
★ 2,2K⑂ 249Python
MPL-2.0Q98
Inferentie, implementatie en runtime
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
★ 2,2K⑂ 211Python
Apache-2.0Q98
Agents en multi-agents
future-agi/future-agi
An open-source, self-hostable platform for evaluating, tracing, simulating, protecting, routing, and optimizing LLM and AI-agent applications.
★ 1,7K⑂ 498Python
Apache-2.0Q98
MCP en toolaanroepen
alpic-ai/skybridge
Skybridge is an open-source, full-stack TypeScript and React framework for building type-safe MCP Apps, MCP servers, and ChatGPT Apps across supported UI-enabled MCP clients.
★ 2K⑂ 133TypeScript
MITQ98
Agents en multi-agents
makecindy/cindy
Cindy is an open-source desktop and mobile AI-agent client that works with local files, signed-in applications, browsers, computers and phones. It supports Claude Code and Codex harnesses, multiple model-access options, persistent workspace context, automation and MCP integrations.
★ 2,1K⑂ 283TypeScript
Apache-2.0Q98
Inferentie, implementatie en runtime
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
★ 1,5K⑂ 279C++
Apache-2.0Q98
Agents en multi-agents
0xlazai/alith
A simple, composable, and high-performance AI agent framework designed for Web3 and Crypto, enabling developers to build, deploy, and manage on-chain AI agents with multi-language support and LazAI Gateway integration.
★ 44⑂ 31Rust
Apache-2.0Q88