A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.