Lời nói, âm thanh & giọng nói

ChatTTS

2noise/ChatTTS

Conversational Chinese and English speech generation with speaker and prosody controls; released model weights are noncommercial.

★ 39,8KSố sao
⑂ 4,2KFork
61Vấn đề đang mở
PythonNgôn ngữ
AGPL-3.0Giấy phép
Q90Điểm biên tập

Tổng quan

ChatTTS emphasizes conversational delivery and provides code, weights, and examples. Evaluate text fidelity and latency alongside naturalness. Code uses AGPLv3+ while released weights use CC BY-NC 4.0; these are different permissions.

Tính năng chính

  • Chinese and English
  • Speaker parameters
  • Prosody controls
  • Python inference
  • WebUI

Yêu cầu, cài đặt và bắt đầu nhanh

Follow repository requirements and model setup; try python examples/web/webui.py. Skip optional components the README marks as experimental or discouraged.

Cách sử dụng

Test numbers, abbreviations, mixed languages, and repeated generations with fixed speaker parameters. Review omissions and audio artifacts before adding controls.

How it works
Text, speaker parameters, and control tokens influence speech generation, including pauses and laughter. Applications handle playback and delivery.

Audience and requirements
Speech researchers and noncommercial prototype teams. Python, model weights, audio dependencies, and suitable compute.

Practical use cases
Noncommercial voice-assistant research; prosody experiments; speech demos.

Limitations and selection
Noncommercial weights. Output can be unstable or inaccurate. Real-time assistants need additional turn-taking and playback infrastructure.

Related projects and selection
gradio-app/gradio:Possible complement: compare text, controls, and audio in an evaluation interface.

ollama/ollama:Pipeline: generate text locally before speech synthesis; orchestration and model licensing remain separate.

Source review
Editorial analysis of upstream sources, without runtime or benchmark testing. Proposed workflows are editorial suggestions.

Khả năng tương thích của mô hình và trường hợp sử dụng

Uses ChatTTS speech weights and speaker parameters; a text LLM can supply upstream responses.

Ghi chú về giấy phép và rủi ro

Code: AGPLv3+. Released model: CC BY-NC 4.0, for noncommercial education and research as described upstream.

Editorial source review 2026-09-09T05:00:00.950Z. README and live repository page verified; current stars/forks from GitHub HTML. Last-push metadata retained from 2026-09-05 discovery snapshot. No runtime benchmark. Integration proposals are editorial analysis.

Phát hành và bảo trì

Reviewed 2026-09-09. Counters come from repository pages; features are based on upstream documentation. See Releases in the source links. Editorial analysis of upstream sources, without runtime or benchmark testing. Proposed workflows are editorial suggestions.