Robótica y Edge AI
microsoft/edgeai-for-beginners
This course is designed to guide beginners through the exciting world of Edge AI, covering fundamental concepts, popular models, inference techniques, device-specific applications, model optimization, and the development of intelligent Edge AI agents.
★ 1,6K⑂ 362Jupyter Notebook
MITQ83
Inferencia, implementación y tiempo de ejecución
PaddlePaddle/Paddle-Lite
PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)
★ 7,3K⑂ 1,6KC++
Apache-2.0Q91
Visión por computadora
facebookresearch/sam2
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 19,6K⑂ 2,5KJupyter Notebook
Apache-2.0Q79
Inferencia, implementación y tiempo de ejecución
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
★ 13,2K⑂ 2,4KC++
Apache-2.0Q94
Inferencia, implementación y tiempo de ejecución
NVIDIA/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
★ 3,3K⑂ 513Python
Apache-2.0Q86
Inferencia, implementación y tiempo de ejecución
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3,7K⑂ 755Python
Apache-2.0Q98
Inferencia, implementación y tiempo de ejecución
theopenco/llmgateway
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
★ 1,6K⑂ 172TypeScript
NOASSERTIONQ98
Inferencia, implementación y tiempo de ejecución
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 32,2K⑂ 8,1KPython
Apache-2.0Q98
Inferencia, implementación y tiempo de ejecución
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 89,5K⑂ 21KPython
Apache-2.0Q98
Ajuste fino, entrenamiento y datos
deepspeedai/deepspeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43K⑂ 4,9KPython
Apache-2.0Q94
Agentes y multiagente
opencsgs/csghub
CSGHub is an open-source, on-premise platform for managing the full lifecycle of Large Language Model assets, including models, datasets, spaces, and code, offering functionality comparable to a private Hugging Face.
★ 4,2K⑂ 525Vue
Apache-2.0Q98
MCP y llamadas a herramientas
lemonade-sdk/lemonade
Lemonade is an open-source local AI server that enables users to run optimized Large Language Models (LLMs), speech, and image generation models directly on their own GPUs and NPUs, providing a free and private alternative to cloud APIs.
★ 5,2K⑂ 436C++
Apache-2.0Q98