FastDeploy
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
ToolAI.io · GitHub চ্যানেল
GitHub-এ সম্প্রতি যোগ করা ওপেন-সোর্স AI প্রকল্প আবিষ্কার করুন—রিপোজিটরির তথ্য, সেটআপ নোট, লাইসেন্স ও সংশ্লিষ্ট রিসোর্সসহ।
পাবলিক রিপোজিটরির তথ্য
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
theopenco/llmgateway
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
deepspeedai/deepspeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
opencsgs/csghub
CSGHub is an open-source, on-premise platform for managing the full lifecycle of Large Language Model assets, including models, datasets, spaces, and code, offering functionality comparable to a private Hugging Face.
lemonade-sdk/lemonade
Lemonade is an open-source local AI server that enables users to run optimized Large Language Models (LLMs), speech, and image generation models directly on their own GPUs and NPUs, providing a free and private alternative to cloud APIs.
katanemo/plano
Plano is an AI-native proxy server and data plane built in Rust that centralizes LLM routing, agent orchestration, observability, and guardrails, allowing developers to focus on core agent logic rather than infrastructure plumbing.
xorbitsai/inference
A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.
huggingface/transformers
A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.
openvinotoolkit/openvino
Open-source toolkit for optimizing and deploying AI inference across edge-to-cloud environments.
kvcache-ai/mooncake
Mooncake is a C++ infrastructure project for large-scale LLM inference and training. It separates prefill, decode, and storage resources while providing high-performance transfer, distributed KV-cache storage, and fault-tolerant expert-parallel communication.
পৃষ্ঠা 2 / 4 · 41টি প্রজেক্ট