evals — Model evaluation framework
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
ToolAI.io · GitHub چینل
LLM، ایجنٹ، MCP، RAG اور کوڈنگ کے موضوعات میں فعال اوپن سورس AI ریپوزٹریز کو اسٹارز اور تصدیق شدہ پروجیکٹ معلومات کے ساتھ ٹریک کریں۔
عوامی ریپوزٹری کا ڈیٹا
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
QwenLM/Qwen
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
zai-org/GLM-4.5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
FoundationAgents/MetaGPT
A multi-agent framework organizing software-development tasks through roles and collaborative workflows.
QwenLM/Qwen3
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud
MiniMax-AI/MiniMax-M2
MiniMax-M2, a model built for Max coding & agentic workflows.
deepseek-ai/DeepSeek-Coder
DeepSeek Coder: Let the Code Write Itself
deepseek-ai/DeepSeek-V3
An open-source project focused on large language model research and inference.
deepseek-ai/DeepSeek-R1
An open-source project focused on reasoning model research and evaluation.
صفحہ 10 / 10 · 118 پروجیکٹس