ToolAI.io · GitHub Channel

AI Project Directory on GitHub

Browse the complete ToolAI directory of open-source AI projects on GitHub, organized by topic, language and license.

Public repository data

Project index

15 projects
Screenshot of promptfoo
Agents & Multi-Agent

promptfoo

promptfoo/promptfoo

An open-source CLI and library for evaluating, testing, and red-teaming LLM applications, RAGs, and agents. It enables side-by-side model comparison and vulnerability scanning with declarative configs and CI/CD integration.

★ 23.9K⑂ 2.2KTypeScript
MITQ98
Agents & Multi-Agent

Opik: Open-Source LLM Observability, Evaluation and Agent Tracing

comet-ml/opik

Opik is an Apache-2.0-licensed platform for tracing, evaluating, debugging and monitoring LLM applications, RAG systems and multi-step agent workflows. It supports self-hosting, a hosted Comet.com option, client SDKs, a REST API, automated evaluations, prompt experimentation and production dashboards.

★ 21.1K⑂ 1.7KPython
Apache-2.0Q98
Evaluation, Observability & Safety

SkillSpector

NVIDIA/SkillSpector

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks.

★ 13.8K⑂ 1.1KPython
Apache-2.0Q98
Screenshot of BISHENG
Agents & Multi-Agent

BISHENG

dataelement/bisheng

BISHENG is an open-source LLM application DevOps platform designed for next-generation enterprise AI applications, offering comprehensive features like GenAI workflow orchestration, RAG, Agent management, and enterprise-grade system controls.

★ 11.8K⑂ 1.9KPython
Apache-2.0Q98
Agents & Multi-Agent

LangWatch

langwatch/langwatch

An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.

★ 3.5K⑂ 345TypeScript
Apache-2.0Q98
Screenshot of TruLens: Evaluation and Tracking for LLM Experiments and AI Agents
Agents & Multi-Agent

TruLens: Evaluation and Tracking for LLM Experiments and AI Agents

truera/trulens

TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.

★ 3.5K⑂ 319Python
MITQ98
Screenshot of EvalScope
RAG & Knowledge Systems

EvalScope

modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

★ 3.2K⑂ 440Python
Apache-2.0Q98
Agents & Multi-Agent

Claude Cookbooks: From API Examples to Evaluatable Business Assistants

anthropics/claude-cookbooks

Anthropic’s collection of Claude development examples covers classification, summarization, retrieval augmentation, tool use, image understanding, and evaluation. It is suitable for learning from specific tasks and adapting the code.

★ 52.6K⑂ 6.3KJupyter Notebook
MITQ92
Evaluation, Observability & Safety

netdata

netdata/netdata

Real-time system and application monitoring with AI-assisted operational analysis.

★ 80.4K⑂ 6.6KGo
GPL-3.0Q90
Fine-tuning, Training & Data

scikit-learn

scikit-learn/scikit-learn

A Python machine learning library for classification, regression, clustering and model evaluation.

★ 67.2K⑂ 27.4KPython
BSD-3-ClauseQ90
Evaluation, Observability & Safety

strix

usestrix/strix

An AI-assisted vulnerability discovery tool for authorized application security testing.

★ 60.6K⑂ 6.6KPython
Apache-2.0Q90

Page 1 / 2 · 15 projects

Recently updated

TypeSafe Python SDKtypesafe-ai/typesafe-sdk-python★ 218 TypeSafe JavaScript SDKtypesafe-ai/typesafe-sdk-js★ 231 TypeSafe Agent Skillstypesafe-ai/skills★ 2K cherry-studioCherryHQ/cherry-studio★ 51.5K onyxonyx-dot-app/onyx★ 32K siyuansiyuan-note/siyuan★ 46.2K

Most starred

ECCaffaan-m/ECC★ 248.8K Hermes Agentnousresearch/hermes-agent★ 227.1K tensorflowtensorflow/tensorflow★ 198.8K AutoGPTSignificant-Gravitas/AutoGPT★ 187.1K ollamaollama/ollama★ 180.2K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174.9K