memra
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
ToolAI.io · GitHub 채널
LLM, 에이전트, MCP, RAG 및 코딩 분야의 활발한 오픈 소스 AI 리포지토리를 추적하고, 스타 수와 검증된 프로젝트 정보를 제공합니다.
공개 리포지토리 데이터
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
taichuy/1flowbase
An open-source AI gateway that allows local agent clients to publish fusion-style multi-model workflows as OpenAI and Claude-compatible virtual models, providing full observability into traces, tokens, latency, and costs.
soapbucket/sbproxy
An open-source, self-hosted Enterprise AI Gateway and LLM proxy written in Rust, providing an OpenAI-compatible API for over 60 providers and supporting local model hosting, MCP tool federation, and robust traffic governance.
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
bionic-gpt/bionic-gpt
Bionic is a Rust-based, on-premise alternative to ChatGPT designed for private generative AI deployments. It provides a familiar chat interface, local or remote model access, team controls, retrieval-augmented assistants, data integrations, observability, and Kubernetes-oriented scaling.
anthropics/skills
Public repository for Agent Skills
anthropics/claude-plugins-official
Official, Anthropic-managed directory of high quality Claude Code Plugins
dotharness/dotcraft
An open-source, self-hosted, project-scoped AI agent runtime that provides persistent sessions, memory, background work, and automations shared across Desktop, CLI, bots, and applications.
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
artokun/comfyui-mcp
A local-first, agent-native control plane for ComfyUI that provides an MCP server and sidebar agent to generate images, video, and audio, author and run workflows, and edit live graphs using natural language across any LLM.
infino-ai/infino
Infino is a fast retrieval engine that executes SQL, full-text search, and vector search over a single copy of data stored natively as Parquet on object storage.
2 / 11페이지 · 프로젝트 130개