evals:Model evaluation framework
openai/evals
ToolAI 对 openai/evals 的客观项目介绍。Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
ToolAI.io · GitHub 频道
一个基于事实的 LLM、智能体、MCP、RAG 和 AI 开发项目索引;每个条目均包含许可证、设置、下载、截图及相关资源。
公开仓库数据
openai/evals
ToolAI 对 openai/evals 的客观项目介绍。Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
QwenLM/Qwen3-Coder
ToolAI 对 QwenLM/Qwen3-Coder 的客观项目介绍。Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
QwenLM/Qwen
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
zai-org/GLM-4.5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
FoundationAgents/MetaGPT
以角色分工和协作流程组织软件开发等任务的多智能体框架。
QwenLM/Qwen3
ToolAI 对 QwenLM/Qwen3 的客观项目介绍。Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
MiniMax-AI/MiniMax-M2
MiniMax-M2, a model built for Max coding & agentic workflows.
deepseek-ai/DeepSeek-Coder
ToolAI 对 deepseek-ai/DeepSeek-Coder 的客观项目介绍。DeepSeek Coder: Let the Code Write Itself
deepseek-ai/DeepSeek-V3
ToolAI 对 deepseek-ai/DeepSeek-V3 的客观项目介绍。An open-source project focused on large language model research and inference.
deepseek-ai/DeepSeek-R1
ToolAI 对 deepseek-ai/DeepSeek-R1 的客观项目介绍。An open-source project focused on reasoning model research and evaluation.
第 10 / 10 页 · 118 个项目