Tencent AI Releases AuK: An MIT-Licensed Open-Source Model for Unified Speech Generation and Editing
Key Info
Tencent AI has released AuK, an MIT-licensed open-source foundation model for unified speech generation and editing, along with AuK-Flash for faster 4-step inference. The model runs on a single consumer GPU and has been supported by SGLang-Omni since day one.
Highlights
- One interface combines natural-language instructions with reference audio, enabling zero-shot TTS, instruction-controlled generation, content editing, whisper conversion, de-accenting, timbre/style/emotion edits, speed/pitch control, enhancement, denoising, multi-speaker and music separation.
- AuK-Flash uses 4-step inference and delivers roughly 4.5× faster generation under matched conditions.
- Code, weights, and related resources are being released under the MIT license, making it practical to run on a single consumer GPU.