evals — Model evaluation framework
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
ToolAI.io · Kênh GitHub
Theo dõi các kho lưu trữ AI mã nguồn mở đang hoạt động trong các chủ đề LLM, tác nhân, MCP, RAG và lập trình, kèm số sao và thông tin dự án đã xác minh.
Dữ liệu kho lưu trữ công khai
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
QwenLM/Qwen
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
zai-org/GLM-4.5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
FoundationAgents/MetaGPT
A multi-agent framework organizing software-development tasks through roles and collaborative workflows.
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
MiniMax-AI/MiniMax-M2
MiniMax-M2, a model built for Max coding & agentic workflows.
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
deepseek-ai/DeepSeek-V3
An open-source project focused on large language model research and inference.
deepseek-ai/DeepSeek-R1
An open-source project focused on reasoning model research and evaluation.
Trang 10 / 10 · 118 dự án