Qwen Audio 3.0 TTS solves the problem of generating high-quality, natural-sounding speech for various applications such as audiobooks, advertising, video dubbing, and podcast production. It also addresses the issue of preserving speaker identity in noisy or reverberant reference audio.
Qwen Audio 3.0 TTS is a production-oriented text-to-speech model that generates natural-sounding speech from text input. It supports free-form instructions for role, emotion, pace, timbre, style, and accent, and allows for inline tags for precise control. The model can generate long-form narration up to 2 minutes without stitching short clips and preserves speaker identity even when reference audio is noisy or reverberant. Qwen Audio 3.0 TTS has two models: Plus for quality-focused production and Flash for low-latency interactive experiences.
Have feedback for the maker?
Sign in to leave a review, report a bug, or suggest a feature.
Wendy Xu
Independent maker exploring practical AI tools and reviewable creative workflows.