A PyTorch-based audio generation research library containing components such as MusicGen, AudioGen, and EnCodec. It supports text-conditioned generation, some melody-conditioned tasks, and model training workflows.
A PaddlePaddle-based speech toolkit covering speech recognition, text-to-speech, punctuation restoration, and other audio tasks, with command-line, Python interface, and service examples.
An open speech model family from Qwen with preset speakers, text-driven voice design and reference-audio cloning for narration and speech applications.