evaluation safety
evals — Model evaluation framework
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
★ 19K⑂ 3KPython
ToolAI.io · GitHub Channel
An objective, fact-checked index of LLM, agent, MCP, RAG and AI development projects. Each entry includes license, setup, download, screenshots and related resources, with human review before publication.
Public repository data
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
NVIDIA/garak
the LLM vulnerability scanner