ToolAI.io · Chaîne GitHub

Nouveaux projets AI sur GitHub

Découvrez les nouveaux projets AI open source ajoutés sur GitHub, avec les informations du dépôt, les instructions de configuration, les licences et les ressources associées.

Données du dépôt public

Index du projet

9 Projets
Capture d’écran de EvalScope
RAG et systèmes de connaissances

EvalScope

modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

★ 3,2K⑂ 440Python
Apache-2.0Q98
Capture d’écran de TruLens: Evaluation and Tracking for LLM Experiments and AI Agents
Agents et multi-agents

TruLens: Evaluation and Tracking for LLM Experiments and AI Agents

truera/trulens

TruLens is an open-source, OpenTelemetry-native evaluation and tracking library for LLM applications and AI agents, enabling developers to trace every step, score performance with LLM judges, and compare app versions.

★ 3,5K⑂ 319Python
MITQ98
Capture d’écran de BISHENG
Agents et multi-agents

BISHENG

dataelement/bisheng

BISHENG is an open-source LLM application DevOps platform designed for next-generation enterprise AI applications, offering comprehensive features like GenAI workflow orchestration, RAG, Agent management, and enterprise-grade system controls.

★ 11,8K⑂ 1,9KPython
Apache-2.0Q98
Capture d’écran de promptfoo
Agents et multi-agents

promptfoo

promptfoo/promptfoo

An open-source CLI and library for evaluating, testing, and red-teaming LLM applications, RAGs, and agents. It enables side-by-side model comparison and vulnerability scanning with declarative configs and CI/CD integration.

★ 23,9K⑂ 2,2KTypeScript
MITQ98
Agents et multi-agents

Opik: Open-Source LLM Observability, Evaluation and Agent Tracing

comet-ml/opik

Opik is an Apache-2.0-licensed platform for tracing, evaluating, debugging and monitoring LLM applications, RAG systems and multi-step agent workflows. It supports self-hosting, a hosted Comet.com option, client SDKs, a REST API, automated evaluations, prompt experimentation and production dashboards.

★ 21,1K⑂ 1,7KPython
Apache-2.0Q98
Agents et multi-agents

LangWatch

langwatch/langwatch

An open-core platform for evaluating, testing, tracing, and monitoring LLM applications and AI agents before release and in production.

★ 3,5K⑂ 345TypeScript
Apache-2.0Q98
Capture d’écran de evals — Model evaluation framework
Évaluation, observabilité et sécurité

evals — Model evaluation framework

openai/evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks

★ 19,2K⑂ 3,1KPython
NOASSERTIONQ82

Récemment mis à jour

codex — Coding agent and developer workflowsopenai/codex★ 106,5K NemoClaw — Agent runtime and deployment toolingNVIDIA/NemoClaw★ 22,2K XERJxerj-org/xerj★ 1,4K LangWatchlangwatch/langwatch★ 3,5K OrchestKityonatangross/orchestkit★ 222 hal0hal0ai/hal0★ 67

Les plus étoilés

Hermes Agentnousresearch/hermes-agent★ 227,1K markitdown — Document conversion and extractionmicrosoft/markitdown★ 174,2K skills — Reusable agent skills and workflowsanthropics/skills★ 169,9K Hugging Face Transformershuggingface/transformers★ 163,3K Firecrawlfirecrawl/firecrawl★ 161,1K LangChainlangchain-ai/langchain★ 143,6K