机器人与边缘 AI
microsoft/edgeai-for-beginners
This course is designed to guide beginners through the exciting world of Edge AI, covering fundamental concepts, popular models, inference techniques, device-specific applications, model optimization, and the development of intelligent Edge AI agents.
★ 1.6K⑂ 362Jupyter Notebook
MITQ83
推理、部署与运行时
PaddlePaddle/Paddle-Lite
PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)
★ 7.3K⑂ 1.6KC++
Apache-2.0Q91
计算机视觉
facebookresearch/sam2
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 19.6K⑂ 2.5KJupyter Notebook
Apache-2.0Q79
推理、部署与运行时
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
★ 13.2K⑂ 2.4KC++
Apache-2.0Q94
推理、部署与运行时
NVIDIA/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
★ 3.3K⑂ 513Python
Apache-2.0Q86
推理、部署与运行时
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7K⑂ 755Python
Apache-2.0Q98
推理、部署与运行时
theopenco/llmgateway
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
★ 1.6K⑂ 172TypeScript
NOASSERTIONQ98
推理、部署与运行时
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 32.2K⑂ 8.1KPython
Apache-2.0Q98
推理、部署与运行时
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 89.5K⑂ 21KPython
Apache-2.0Q98
微调、训练与数据
deepspeedai/deepspeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43K⑂ 4.9KPython
Apache-2.0Q94
智能体与多智能体
opencsgs/csghub
CSGHub 是一个开源的本地化部署平台,用于管理大语言模型资产的全生命周期,包括模型、数据集、空间和代码,提供类似于私有化 Hugging Face 的功能。
★ 4.2K⑂ 525Vue
Apache-2.0Q98
MCP 与工具调用
lemonade-sdk/lemonade
Lemonade 是一个开源的本地 AI 服务器,允许用户直接在自己的 GPU 和 NPU 上运行优化的大语言模型(LLM)、语音及图像生成模型,提供了一种免费且私有的云 API 替代方案。
★ 5.2K⑂ 436C++
Apache-2.0Q98