Meta Launches Muse Voice Transcribe: Real-Time Audio Perception Model with Streaming ASR and Diarization

AI at Meta ·

Key Info

Meta Superintelligence Labs has launched Muse Voice Transcribe, a real-time audio perception model that performs streaming speech-to-text, speaker diarization, and endpointing in a single model, claiming state-of-the-art streaming ASR performance.

Highlights

  • Available through the Meta Model API, Meta AI for Mac, and Muse Code.
  • Uses adaptive delay to reach the Pareto frontier in the speed-accuracy trade-off for time-to-final transcription.
  • Natively supports diarization for 20+ speakers, simplifying multi-speaker transcription workflows.
Loading...