로보틱스 및 엣지 AI
microsoft/edgeai-for-beginners
This course is designed to guide beginners through the exciting world of Edge AI, covering fundamental concepts, popular models, inference techniques, device-specific applications, model optimization, and the development of intelligent Edge AI agents.
★ 1.6K⑂ 362Jupyter Notebook
MITQ83
추론, 배포 및 런타임
PaddlePaddle/Paddle-Lite
PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)
★ 7.3K⑂ 1.6KC++
Apache-2.0Q91
컴퓨터 비전
facebookresearch/sam2
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 19.6K⑂ 2.5KJupyter Notebook
Apache-2.0Q79
추론, 배포 및 런타임
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
★ 13.2K⑂ 2.4KC++
Apache-2.0Q94
추론, 배포 및 런타임
NVIDIA/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
★ 3.3K⑂ 513Python
Apache-2.0Q86
추론, 배포 및 런타임
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7K⑂ 755Python
Apache-2.0Q98
추론, 배포 및 런타임
theopenco/llmgateway
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
★ 1.6K⑂ 172TypeScript
NOASSERTIONQ98
추론, 배포 및 런타임
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 32.2K⑂ 8.1KPython
Apache-2.0Q98
추론, 배포 및 런타임
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 89.5K⑂ 21KPython
Apache-2.0Q98
파인튜닝, 학습 및 데이터
deepspeedai/deepspeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43K⑂ 4.9KPython
Apache-2.0Q94
에이전트 및 멀티 에이전트
opencsgs/csghub
CSGHub is an open-source, on-premise platform for managing the full lifecycle of Large Language Model assets, including models, datasets, spaces, and code, offering functionality comparable to a private Hugging Face.
★ 4.2K⑂ 525Vue
Apache-2.0Q98
MCP 및 도구 호출
lemonade-sdk/lemonade
Lemonade is an open-source local AI server that enables users to run optimized Large Language Models (LLMs), speech, and image generation models directly on their own GPUs and NPUs, providing a free and private alternative to cloud APIs.
★ 5.2K⑂ 436C++
Apache-2.0Q98