Project screenshots
Overview
Xinference simplifies the deployment and serving of state-of-the-art built-in or custom AI models using a single command. It supports distributed inference across multiple devices or machines and integrates with various inference engines like vLLM and ggml. The framework provides an OpenAI-compatible RESTful API, enabling seamless migration and interoperability for applications built on the OpenAI API standard.
Key features
- Model Serving Made Easy: Deploy models with a single command
- State-of-the-Art Built-in Models: Access cutting-edge open-source models effortlessly
- Heterogeneous Hardware Utilization: Intelligently use GPUs and CPUs via ggml
- Flexible API and Interfaces: OpenAI-compatible RESTful API, RPC, CLI, and WebUI
- Distributed Deployment: Distribute inference across multiple devices or machines
- Built-in Integration with Third-Party Libraries: LangChain, LlamaIndex, Dify, Chatbox, Xagent
- Auto batching for improved throughput
- Agent-native Serving via Xagent integration
Requirements, installation and quick start
Usage
Model compatibility and use cases
Supports a wide range of built-in models including Llama3, ChatGLM, GLM4, Flan-T5, Gemma, Mistral, Qwen, Whisper, WizardLM, MiniMax-M3, VibeThinker, Nex-N2, Unlimited-OCR, Ornith-1.0-35B, MiniCPM5-1B, jina-embeddings-v5, and MiniCPM-V-4.6. Also supports custom models.
License and risk notes
Licensed under Apache-2.0.
Editorial verification 2026-08-09: repository URL, owner, description, license and repository statistics were reviewed. License metadata: Apache-2.0. README was fetched for the channel draft; re-check repository dependencies, releases and model terms before production use.
Release and maintenance
Xinference 3.0.0 is available with migration notes and breaking changes. Enhancements include Agent-native Serving with Xagent, Auto batch for concurrent requests, Xllamacpp for continuous batching, distributed inference across workers, and VLLM shared KV cache across multiple replicas.