evals — Model evaluation framework
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
ToolAI.io · Chaîne GitHub
Suivez les dépôts AI open source actifs dans les domaines des LLM, des agents, de MCP, de RAG et du code, avec leur nombre d’étoiles et des informations vérifiées sur les projets.
Données du dépôt public
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
QwenLM/Qwen
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
zai-org/GLM-4.5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
FoundationAgents/MetaGPT
A multi-agent framework organizing software-development tasks through roles and collaborative workflows.
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
MiniMax-AI/MiniMax-M2
MiniMax-M2, a model built for Max coding & agentic workflows.
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
deepseek-ai/DeepSeek-V3
An open-source project focused on large language model research and inference.
deepseek-ai/DeepSeek-R1
An open-source project focused on reasoning model research and evaluation.
Page 10 / 10 · 118 projets