Agents et multi-agents
infiniflow/ragflow
RAGFlow is an open-source Retrieval-Augmented Generation engine that combines document ingestion, retrieval, grounded citations, configurable language and embedding models, and agent capabilities to provide context for LLM applications.
★ 86,7K⑂ 10,2KGo
Apache-2.0Q98
Inférence, déploiement et exécution
kvcache-ai/mooncake
Mooncake is a C++ infrastructure project for large-scale LLM inference and training. It separates prefill, decode, and storage resources while providing high-performance transfer, distributed KV-cache storage, and fault-tolerant expert-parallel communication.
★ 6,1K⑂ 1KC++
Apache-2.0Q98
Agents et multi-agents
mlflow/mlflow
An open-source AI engineering platform for tracing, evaluating, monitoring, optimizing, governing, and deploying agents, LLM applications, and machine-learning models.
★ 27,3K⑂ 6,1KPython
Apache-2.0Q98
Réglage fin, entraînement et données
langfuse/langfuse
An open-source LLM engineering platform for tracing, evaluating, debugging, and improving AI applications, with prompt management, datasets, metrics, a playground, APIs, and managed or self-hosted deployment options.
★ 32,4K⑂ 3,5KTypeScript
NOASSERTIONQ98
Inférence, déploiement et exécution
llm-d/llm-d
llm-d is an open-source serving stack that adds distributed orchestration, routing, cache management, autoscaling, and batch processing around model servers such as vLLM and SGLang. It targets high-scale production inference on Kubernetes and modern hardware accelerators.
★ 4K⑂ 649Shell
Apache-2.0Q98
Agents et multi-agents
alibaba/open-code-review
An open-source AI code review CLI from Alibaba that combines deterministic review pipelines with an LLM agent to produce structured, line-level feedback.
★ 18K⑂ 1,2KGo
Apache-2.0Q98
Agents et multi-agents
giskard-ai/giskard-oss
An open-source Python library for evaluating and testing LLM-based and agentic systems, including multi-turn evaluations, LLM-as-judge checks and automated vulnerability scanning.
★ 5,7K⑂ 513Python
Apache-2.0Q98
MCP et appels d’outils
maximhq/bifrost
Bifrost is a Go-based AI gateway that provides a unified, OpenAI-compatible API for 23+ AI providers. It supports routing, fallbacks, load balancing, semantic caching, governance, observability, multimodal requests, plugins, and Model Context Protocol integrations.
★ 7K⑂ 984Go
Apache-2.0Q98
Réglage fin, entraînement et données
oumi-ai/oumi
Oumi is a Python-based platform for preparing data, training and fine-tuning open-weight foundation models, evaluating results, running inference, and deploying models. It provides configuration recipes and a consistent CLI for local, cluster, and cloud workflows.
★ 9,4K⑂ 784Python
Apache-2.0Q98
Inférence, déploiement et exécution
gpustack/gpustack
GPUStack is an open-source GPU cluster manager for serving AI models and provisioning on-demand, SSH-accessible GPU instances. It orchestrates inference engines including vLLM, SGLang, and TensorRT-LLM across on-premises, Kubernetes, and cloud environments.
★ 5,4K⑂ 604Python
Apache-2.0Q98
MCP et appels d’outils
mempalace/mempalace
A local-first, open-source AI memory system that stores conversation history and project files as verbatim text for semantic retrieval, achieving 96.6% R@5 on LongMemEval without requiring LLMs or API calls.
★ 58K⑂ 7,5KPython
MITQ98
Agents et multi-agents
neuml/txtai
txtai is a Python-based AI framework for semantic and vector search, retrieval-augmented generation, LLM orchestration, autonomous agents and multimodal language-model workflows.
★ 12,8K⑂ 851Python
Apache-2.0Q98