Inferentie, implementatie en runtime
gpustack/gpustack
GPUStack is an open-source GPU cluster manager for serving AI models and provisioning on-demand, SSH-accessible GPU instances. It orchestrates inference engines including vLLM, SGLang, and TensorRT-LLM across on-premises, Kubernetes, and cloud environments.
★ 5,4K⑂ 604Python
Apache-2.0Q98
MCP en toolaanroepen
lemonade-sdk/lemonade
Lemonade is an open-source local AI server that enables users to run optimized Large Language Models (LLMs), speech, and image generation models directly on their own GPUs and NPUs, providing a free and private alternative to cloud APIs.
★ 5,2K⑂ 436C++
Apache-2.0Q98
Agents en multi-agents
embabel/embabel-agent
Embabel is an open-source JVM framework for building strongly typed agentic applications that combine LLM interactions, regular code, tools, and domain models. It supports dynamic planning and can be used from Kotlin or Java.
★ 4,3K⑂ 417Kotlin
Apache-2.0Q98
MCP en toolaanroepen
homeassistant-ai/ha-mcp
A comprehensive Model Context Protocol (MCP) server enabling AI assistants to interact with, configure, build, and debug Home Assistant smart home setups using natural language.
★ 4,3K⑂ 177Python
MITQ98
Agents en multi-agents
opencsgs/csghub
CSGHub is an open-source, on-premise platform for managing the full lifecycle of Large Language Model assets, including models, datasets, spaces, and code, offering functionality comparable to a private Hugging Face.
★ 4,2K⑂ 525Vue
Apache-2.0Q98
MCP en toolaanroepen
oomol-lab/open-connector
An open-source authentication gateway that connects over 1,000 SaaS providers to AI agents through SDK, CLI, MCP, HTTP, and OpenAPI interfaces.
★ 4,1K⑂ 311TypeScript
Apache-2.0Q98
Agents en multi-agents
openagents-org/openagents
An open-source platform providing a collaborative operating system for AI agents, enabling unified workspace management, multi-agent coordination, and network integration without vendor lock-in.
★ 4K⑂ 403TypeScript
Apache-2.0Q98
Inferentie, implementatie en runtime
llm-d/llm-d
llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.
★ 4K⑂ 649Shell
Apache-2.0Q98
MCP en toolaanroepen
atmosphere/atmosphere
Atmosphere is a Java-based real-time framework for building, streaming, and governing AI agents. It provides a unified broadcaster pipeline for WebSocket, SSE, long-polling, and gRPC, alongside built-in governance, human-in-the-loop workflows, and multi-protocol support including MCP, A2A, and AG-UI.
★ 3,8K⑂ 760Java
Apache-2.0Q98
Inferentie, implementatie en runtime
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3,7K⑂ 755Python
Apache-2.0Q98
Agents en multi-agents
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
★ 3,5K⑂ 345TypeScript
Apache-2.0Q98
Agents en multi-agents
truera/trulens
TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.
★ 3,5K⑂ 319Python
MITQ98