Vue d’ensemble
Qwen3-TTS offers separate checkpoints for distinct speech workflows. The 1.7B VoiceDesign model creates voices from natural-language descriptions, CustomVoice provides preset speakers, and Base models accept reference audio for cloning. Choose the workflow before downloading weights: the model families expose different generation interfaces. Useful applications include product narration, course audio and interactive speech features with repeatable speaker settings.
Fonctionnalités clés
- Supports ten languages, including Chinese, English, Japanese and Korean.
- Offers 0.6B and 1.7B model sizes.
- CustomVoice includes nine preset speakers.
- The 1.7B VoiceDesign checkpoint accepts voice descriptions.
- Base supports reference-audio cloning and reusable voice prompts.
- Includes Python APIs, a web demo and streaming generation capabilities.
Prérequis, installation et démarrage rapide
2. Run pip install -U qwen-tts.
3. Select a checkpoint from the official Qwen collection and allow storage for its weights.
4. Load it with Qwen3TTSModel using a device and precision appropriate to your hardware.
5. Match generate_custom_voice, generate_voice_design or generate_voice_clone to the checkpoint.
6. Generate and save a short sentence first, checking language, speaker and sample rate before processing longer scripts.
Utilisation
Implementation notes
Reference noise, accent and speaking style affect cloning. Review pauses between long-text segments. Measure first-audio latency and concurrency on the actual deployment hardware instead of treating upstream demonstrations as a service guarantee.
Compatibilité des modèles et cas d’usage
Requires Python, downloaded model weights and suitable compute. GPU and attention-backend compatibility depend on the environment. The 0.6B CustomVoice model does not offer the same instruction control as its 1.7B counterpart.
Licence et notes sur les risques
The repository uses Apache-2.0. Check the selected checkpoint’s model card as well, and obtain appropriate rights to reference recordings.