Parole, audio et voix

ChatTTS

2noise/ChatTTS

Conversational Chinese and English speech generation with speaker and prosody controls; released model weights are noncommercial.

★ 39,8KÉtoiles
⑂ 4,2KForks
61Problèmes ouverts
PythonLangue
AGPL-3.0Licence
Q90Score éditorial

Vue d’ensemble

ChatTTS emphasizes conversational delivery and provides code, weights, and examples. Evaluate text fidelity and latency alongside naturalness. Code uses AGPLv3+ while released weights use CC BY-NC 4.0; these are different permissions.

Fonctionnalités clés

  • Chinese and English
  • Speaker parameters
  • Prosody controls
  • Python inference
  • WebUI

Prérequis, installation et démarrage rapide

Follow repository requirements and model setup; try python examples/web/webui.py. Skip optional components the README marks as experimental or discouraged.

Utilisation

Test numbers, abbreviations, mixed languages, and repeated generations with fixed speaker parameters. Review omissions and audio artifacts before adding controls.

How it works
Text, speaker parameters, and control tokens influence speech generation, including pauses and laughter. Applications handle playback and delivery.

Audience and requirements
Speech researchers and noncommercial prototype teams. Python, model weights, audio dependencies, and suitable compute.

Practical use cases
Noncommercial voice-assistant research; prosody experiments; speech demos.

Limitations and selection
Noncommercial weights. Output can be unstable or inaccurate. Real-time assistants need additional turn-taking and playback infrastructure.

Related projects and selection
gradio-app/gradio:Possible complement: compare text, controls, and audio in an evaluation interface.

ollama/ollama:Pipeline: generate text locally before speech synthesis; orchestration and model licensing remain separate.

Source review
Editorial analysis of upstream sources, without runtime or benchmark testing. Proposed workflows are editorial suggestions.

Compatibilité des modèles et cas d’usage

Uses ChatTTS speech weights and speaker parameters; a text LLM can supply upstream responses.

Licence et notes sur les risques

Code: AGPLv3+. Released model: CC BY-NC 4.0, for noncommercial education and research as described upstream.

Editorial source review 2026-09-09T05:00:00.950Z. README and live repository page verified; current stars/forks from GitHub HTML. Last-push metadata retained from 2026-09-05 discovery snapshot. No runtime benchmark. Integration proposals are editorial analysis.

Publication et maintenance

Reviewed 2026-09-09. Counters come from repository pages; features are based on upstream documentation. See Releases in the source links. Editorial analysis of upstream sources, without runtime or benchmark testing. Proposed workflows are editorial suggestions.