Tổng quan
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Yêu cầu, cài đặt và bắt đầu nhanh
Cách sử dụng
Khả năng tương thích của mô hình và trường hợp sử dụng
Khả năng tương thích của mô hình không được nêu trong siêu dữ liệu của kho lưu trữ.
Ghi chú về giấy phép và rủi ro
Apache-2.0
ToolAI metadata-only listing: repository metadata is public; editorial content and manual verification are still pending. Quality score was recomputed from live GitHub repository facts on 2026-08-22 UTC using the seven-dimension channel formula.