프로젝트 스크린샷
개요
Xinference simplifies the deployment and serving of state-of-the-art built-in or custom AI models using a single command. It supports distributed inference across multiple devices or machines and integrates with various inference engines like vLLM and ggml. The framework provides an OpenAI-compatible RESTful API, enabling seamless migration and interoperability for applications built on the OpenAI API standard.
주요 기능
- Model Serving Made Easy: Deploy models with a single command
- State-of-the-Art Built-in Models: Access cutting-edge open-source models effortlessly
- Heterogeneous Hardware Utilization: Intelligently use GPUs and CPUs via ggml
- Flexible API and Interfaces: OpenAI-compatible RESTful API, RPC, CLI, and WebUI
- Distributed Deployment: Distribute inference across multiple devices or machines
- Built-in Integration with Third-Party Libraries: LangChain, LlamaIndex, Dify, Chatbox, Xagent
- Auto batching for improved throughput
- Agent-native Serving via Xagent integration
요구 사항, 설치 및 빠른 시작
사용 정보
모델 호환성 및 사용 사례
Supports a wide range of built-in models including Llama3, ChatGLM, GLM4, Flan-T5, Gemma, Mistral, Qwen, Whisper, WizardLM, MiniMax-M3, VibeThinker, Nex-N2, Unlimited-OCR, Ornith-1.0-35B, MiniCPM5-1B, jina-embeddings-v5, and MiniCPM-V-4.6. Also supports custom models.
라이선스 및 위험 참고 사항
Licensed under Apache-2.0.
Editorial verification 2026-08-09: repository URL, owner, description, license and repository statistics were reviewed. License metadata: Apache-2.0. README was fetched for the channel draft; re-check repository dependencies, releases and model terms before production use.
릴리스 및 유지 관리
Xinference 3.0.0 is available with migration notes and breaking changes. Enhancements include Agent-native Serving with Xagent, Auto batch for concurrent requests, Xllamacpp for continuous batching, distributed inference across workers, and VLLM shared KV cache across multiple replicas.