evals — Model evaluation framework
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
ToolAI.io · Kênh GitHub
Chỉ mục dựa trên dữ kiện về các dự án phát triển LLM, tác nhân, MCP, RAG và AI, kèm giấy phép, thiết lập, tải xuống, ảnh chụp màn hình và tài nguyên liên quan cho từng mục.
Dữ liệu kho lưu trữ công khai
openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks
QwenLM/Qwen3-Coder
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way
Trang 7 / 7 · 75 dự án