Inferencia, implementación y tiempo de ejecución
llm-d/llm-d
llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.
★ 4K⑂ 649Shell
Apache-2.0Q98
Agentes y multiagente
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
★ 3,5K⑂ 345TypeScript
Apache-2.0Q98
Inferencia, implementación y tiempo de ejecución
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
★ 2,7K⑂ 2,1KC++
Apache-2.0Q98
LLM y modelos fundacionales
bionic-gpt/bionic-gpt
Bionic is a Rust-based, on-premise alternative to ChatGPT designed for private generative AI deployments. It provides a familiar chat interface, local or remote model access, team controls, retrieval-augmented assistants, data integrations, observability, and Kubernetes-oriented scaling.
★ 2,4K⑂ 239Rust
NOASSERTIONQ98
Agentes y multiagente
embabel/embabel-agent
Embabel is an open-source JVM framework for building strongly typed agentic applications that combine LLM interactions, regular code, tools, and domain models. It supports dynamic planning and can be used from Kotlin or Java.
★ 4,3K⑂ 417Kotlin
Apache-2.0Q98
Agentes y multiagente
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
★ 2,2K⑂ 249Python
MPL-2.0Q98
Inferencia, implementación y tiempo de ejecución
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
★ 2,2K⑂ 211Python
Apache-2.0Q98
Agentes y multiagente
future-agi/future-agi
An open-source, self-hostable platform for evaluating, tracing, simulating, protecting, routing, and optimizing LLM and AI-agent applications.
★ 1,7K⑂ 498Python
Apache-2.0Q98
MCP y llamadas a herramientas
alpic-ai/skybridge
Skybridge is an open-source, full-stack TypeScript and React framework for building type-safe MCP Apps, MCP servers, and ChatGPT Apps across supported UI-enabled MCP clients.
★ 2K⑂ 133TypeScript
MITQ98
Agentes y multiagente
makecindy/cindy
Cindy is an open-source desktop and mobile AI-agent client that works with local files, signed-in applications, browsers, computers and phones. It supports Claude Code and Codex harnesses, multiple model-access options, persistent workspace context, automation and MCP integrations.
★ 2,1K⑂ 283TypeScript
Apache-2.0Q98
Inferencia, implementación y tiempo de ejecución
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
★ 1,5K⑂ 279C++
Apache-2.0Q98
Agentes y multiagente
0xlazai/alith
A simple, composable, and high-performance AI agent framework designed for Web3 and Crypto, enabling developers to build, deploy, and manage on-chain AI agents with multi-language support and LazAI Gateway integration.
★ 44⑂ 31Rust
Apache-2.0Q88