EvalScope
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
ToolAI.io · Chaîne GitHub
Découvrez les nouveaux projets AI open source ajoutés sur GitHub, avec les informations du dépôt, les instructions de configuration, les licences et les ressources associées.
Données du dépôt public
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
truera/trulens
TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.
dataelement/bisheng
BISHENG is an open-source LLM application DevOps platform designed for next-generation enterprise AI applications, offering comprehensive features like GenAI workflow orchestration, RAG, Agent management, and enterprise-grade system controls.
promptfoo/promptfoo
An open-source CLI and library for evaluating, testing, and red-teaming LLM applications, RAGs, and agents. It enables side-by-side model comparison and vulnerability scanning with declarative configs and CI/CD integration.
comet-ml/opik
Opik is an Apache-2.0-licensed platform for tracing, evaluating, debugging and monitoring LLM applications, RAG systems and multi-step agent workflows. It supports self-hosting, a hosted Comet.com option, client SDKs, a REST API, automated evaluations, prompt experimentation and production dashboards.
giskard-ai/giskard-oss
An open-source Python library for evaluating and testing LLM-based and agentic systems, including multi-turn evaluations, LLM-as-judge checks and automated vulnerability scanning.
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
NVIDIA/garak
the LLM vulnerability scanner
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks