evals — Model evaluation framework
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
ToolAI.io · Canal do GitHub
Acompanhe repositórios de AI de código aberto ativos nos temas de LLM, agentes, MCP, RAG e programação, com número de estrelas e dados de projetos verificados.
Dados do repositório público
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
QwenLM/Qwen
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
zai-org/GLM-4.5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
FoundationAgents/MetaGPT
A multi-agent framework organizing software-development tasks through roles and collaborative workflows.
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
MiniMax-AI/MiniMax-M2
MiniMax-M2, a model built for Max coding & agentic workflows.
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
deepseek-ai/DeepSeek-V3
An open-source project focused on large language model research and inference.
deepseek-ai/DeepSeek-R1
An open-source project focused on reasoning model research and evaluation.
Página 10 / 10 · 118 projetos