LangWatch
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
ToolAI.io · Kênh GitHub
Theo dõi các kho lưu trữ AI mã nguồn mở đang hoạt động trong các chủ đề LLM, tác nhân, MCP, RAG và lập trình, kèm số sao và thông tin dự án đã xác minh.
Dữ liệu kho lưu trữ công khai
langwatch/langwatch
An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.
NVIDIA/garak
the LLM vulnerability scanner
truera/trulens
TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
dataelement/bisheng
BISHENG is an open-source LLM application DevOps platform designed for next-generation enterprise AI applications, offering comprehensive features like GenAI workflow orchestration, RAG, Agent management, and enterprise-grade system controls.
promptfoo/promptfoo
An open-source CLI and library for evaluating, testing, and red-teaming LLM applications, RAGs, and agents. It enables side-by-side model comparison and vulnerability scanning with declarative configs and CI/CD integration.
comet-ml/opik
Opik is an Apache-2.0-licensed platform for tracing, evaluating, debugging and monitoring LLM applications, RAG systems and multi-step agent workflows. It supports self-hosting, a hosted Comet.com option, client SDKs, a REST API, automated evaluations, prompt experimentation and production dashboards.
giskard-ai/giskard-oss
An open-source Python library for evaluating and testing LLM-based and agentic systems, including multi-turn evaluations, LLM-as-judge checks and automated vulnerability scanning.
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks