अवलोकन
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
आवश्यकताएँ, इंस्टॉलेशन और त्वरित शुरुआत
उपयोग
मॉडल संगतता और उपयोग के मामले
रिपॉज़िटरी मेटाडेटा में मॉडल संगतता का उल्लेख नहीं है।
लाइसेंस और जोखिम संबंधी टिप्पणियाँ
Apache-2.0
ToolAI metadata-only listing: repository metadata is public; editorial content and manual verification are still pending. Quality score was recomputed from live GitHub repository facts on 2026-08-22 UTC using the seven-dimension channel formula.