Agents et multi-agents
promptfoo/promptfoo
An open-source CLI and library for evaluating, testing, and red-teaming LLM applications, RAGs, and agents. It enables side-by-side model comparison and vulnerability scanning with declarative configs and CI/CD integration.
★ 23,9K⑂ 2,2KTypeScript
MITQ98
Agents et multi-agents
comet-ml/opik
Opik is an Apache-2.0-licensed platform for tracing, evaluating, debugging and monitoring LLM applications, RAG systems and multi-step agent workflows. It supports self-hosting, a hosted Comet.com option, client SDKs, a REST API, automated evaluations, prompt experimentation and production dashboards.
★ 21,1K⑂ 1,7KPython
Apache-2.0Q98
Agents et multi-agents
dataelement/bisheng
BISHENG is an open-source LLM application DevOps platform designed for next-generation enterprise AI applications, offering comprehensive features like GenAI workflow orchestration, RAG, Agent management, and enterprise-grade system controls.
★ 11,8K⑂ 1,9KPython
Apache-2.0Q98
Agents et multi-agents
giskard-ai/giskard-oss
An open-source Python library for evaluating and testing LLM-based and agentic systems, including multi-turn evaluations, LLM-as-judge checks and automated vulnerability scanning.
★ 5,7K⑂ 513Python
Apache-2.0Q98
Agents et multi-agents
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
★ 3,5K⑂ 345TypeScript
Apache-2.0Q98
Agents et multi-agents
truera/trulens
TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.
★ 3,5K⑂ 319Python
MITQ98
RAG et systèmes de connaissances
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
★ 3,2K⑂ 440Python
Apache-2.0Q98
Évaluation, observabilité et sécurité
NVIDIA/garak
the LLM vulnerability scanner
★ 8,8K⑂ 1,2KPython
Apache-2.0Q84
Évaluation, observabilité et sécurité
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
★ 19,2K⑂ 3,1KPython
NOASSERTIONQ82