Tencent Releases AuK: Open-Source Speech Generation and Editing Model

Tencent Hy ·

Key Info

Tencent has officially released AuK, an open-source foundation model for unified speech generation and editing. It combines natural-language instructions with reference audio through a single interface.

Highlights

  • AuK supports zero-shot TTS, instruction-controlled generation, content editing, whisper conversion, de-accenting, timbre/style/emotion editing, and speed/pitch control.
  • It also provides enhancement, denoising, multi-speaker, and music separation capabilities.
  • AuK-Flash, a faster variant, delivers 4-step inference and runs ~4.5× faster under matched conditions; code, weights, and related resources are being released.
Loading...