Ucapan, audio & suara

Qwen3-TTS: multilingual speech, voice design and cloning

QwenLM/Qwen3-TTS

An open speech model family from Qwen with preset speakers, text-driven voice design and reference-audio cloning for narration and speech applications.

★ 13,4KBintang
⑂ 1,7KFork
57Isu terbuka
PythonBahasa
Apache-2.0Lisensi
Q84Skor editorial

Ringkasan

Qwen3-TTS offers separate checkpoints for distinct speech workflows. The 1.7B VoiceDesign model creates voices from natural-language descriptions, CustomVoice provides preset speakers, and Base models accept reference audio for cloning. Choose the workflow before downloading weights: the model families expose different generation interfaces. Useful applications include product narration, course audio and interactive speech features with repeatable speaker settings.

Fitur utama

  • Supports ten languages, including Chinese, English, Japanese and Korean.
  • Offers 0.6B and 1.7B model sizes.
  • CustomVoice includes nine preset speakers.
  • The 1.7B VoiceDesign checkpoint accepts voice descriptions.
  • Base supports reference-audio cloning and reusable voice prompts.
  • Includes Python APIs, a web demo and streaming generation capabilities.

Persyaratan, instalasi, dan mulai cepat

1. Create an isolated environment; the upstream guide recommends Python 3.12.
2. Run pip install -U qwen-tts.
3. Select a checkpoint from the official Qwen collection and allow storage for its weights.
4. Load it with Qwen3TTSModel using a device and precision appropriate to your hardware.
5. Match generate_custom_voice, generate_voice_design or generate_voice_clone to the checkpoint.
6. Generate and save a short sentence first, checking language, speaker and sample rate before processing longer scripts.

Penggunaan

For product narration, choose one preset speaker, split scripts at natural pauses and review names and technical terms before editing audio into the video. For cloning, use a clean reference recording that you have permission to use and supply its transcript as required by the API. Reuse the same voice prompt across segments, and retain the text, model identifier and generation settings so revised passages can be regenerated.

Implementation notes
Reference noise, accent and speaking style affect cloning. Review pauses between long-text segments. Measure first-audio latency and concurrency on the actual deployment hardware instead of treating upstream demonstrations as a service guarantee.

Kompatibilitas model dan kasus penggunaan

Requires Python, downloaded model weights and suitable compute. GPU and attention-backend compatibility depend on the environment. The 0.6B CustomVoice model does not offer the same instruction control as its 1.7B counterpart.

Catatan lisensi dan risiko

The repository uses Apache-2.0. Check the selected checkpoint’s model card as well, and obtain appropriate rights to reference recordings.

ChatTTS

2noise/ChatTTS

★ 39,8KPython