ToolAI.io · GitHubチャンネル

GitHub上のオープンソースAIプロジェクト

ライセンス、セットアップ、ダウンロード、スクリーンショット、関連リソースを各エントリーに掲載した、LLM、エージェント、MCP、RAG、AI開発プロジェクトの事実に基づくインデックス。

公開リポジトリのデータ

プロジェクトインデックス

9 プロジェクト
promptfooのスクリーンショット
エージェントとマルチエージェント

promptfoo

promptfoo/promptfoo

An open-source CLI and library for evaluating, testing, and red-teaming LLM applications, RAGs, and agents. It enables side-by-side model comparison and vulnerability scanning with declarative configs and CI/CD integration.

★ 23.9K⑂ 2.2KTypeScript
MITQ98
エージェントとマルチエージェント

Opik: Open-Source LLM Observability, Evaluation and Agent Tracing

comet-ml/opik

Opik is an Apache-2.0-licensed platform for tracing, evaluating, debugging and monitoring LLM applications, RAG systems and multi-step agent workflows. It supports self-hosting, a hosted Comet.com option, client SDKs, a REST API, automated evaluations, prompt experimentation and production dashboards.

★ 21.1K⑂ 1.7KPython
Apache-2.0Q98
BISHENGのスクリーンショット
エージェントとマルチエージェント

BISHENG

dataelement/bisheng

BISHENG is an open-source LLM application DevOps platform designed for next-generation enterprise AI applications, offering comprehensive features like GenAI workflow orchestration, RAG, Agent management, and enterprise-grade system controls.

★ 11.8K⑂ 1.9KPython
Apache-2.0Q98
エージェントとマルチエージェント

Giskard OSS: Evaluation, Red Teaming and Test Generation for AI Agents

giskard-ai/giskard-oss

An open-source Python library for evaluating and testing LLM-based and agentic systems, including multi-turn evaluations, LLM-as-judge checks and automated vulnerability scanning.

★ 5.7K⑂ 513Python
Apache-2.0Q98
エージェントとマルチエージェント

LangWatch

langwatch/langwatch

An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.

★ 3.5K⑂ 345TypeScript
Apache-2.0Q98
TruLens: Evaluation and Tracking for LLM Experiments and AI Agentsのスクリーンショット
エージェントとマルチエージェント

TruLens: Evaluation and Tracking for LLM Experiments and AI Agents

truera/trulens

TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.

★ 3.5K⑂ 319Python
MITQ98
EvalScopeのスクリーンショット
RAGとナレッジシステム

EvalScope

modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

★ 3.2K⑂ 440Python
Apache-2.0Q98
evals — Model evaluation frameworkのスクリーンショット
評価・オブザーバビリティ・安全性

evals — Model evaluation framework

openai/evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks

★ 19.2K⑂ 3.1KPython
NOASSERTIONQ82

最近更新

codex — Coding agent and developer workflowsopenai/codex★ 106.5K NemoClaw — Agent runtime and deployment toolingNVIDIA/NemoClaw★ 22.2K XERJxerj-org/xerj★ 1.4K LangWatchlangwatch/langwatch★ 3.5K OrchestKityonatangross/orchestkit★ 222 hal0hal0ai/hal0★ 67

スター数が多い

Hermes Agentnousresearch/hermes-agent★ 227.1K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.2K skills — Reusable agent skills and workflowsanthropics/skills★ 169.9K Hugging Face Transformershuggingface/transformers★ 163.3K Firecrawlfirecrawl/firecrawl★ 161.1K LangChainlangchain-ai/langchain★ 143.6K