Meta launches real-time transcription model for multilingual group conversations

← Back to the feed

Meta launches real-time transcription model for multilingual group conversations

Engadget · 3 hours ago

Meta has launched Muse Voice Transcribe, its first real-time audio transcription model, built by the Meta Superintelligence Lab (MSL). The model can distinguish between more than 20 individual speakers in a single session and switch seamlessly between languages, including mid-sentence "code-switching," making it one of the more advanced speech-to-text tools released so far. The launch comes just days after Google unveiled its own rival system, Gemini 3.5 Transcribe, intensifying competition among major tech firms in real-time audio AI.

Meta says the model was trained on over 70 languages, with 25 validated at launch, and uses an "adaptive delay" system that pauses longer on difficult words while committing quickly to easier ones, aiming to improve accuracy on messy, real-world audio over hour-long sessions. It is already powering dictation features in Meta's new AI Mac app and is available to developers via Muse Code and Meta's Model API, priced at $3 per 1,000 minutes of audio, with a demo also available on Meta's research blog. Unlike Google, which is integrating its equivalent model into Android and Chrome, Meta has not yet said whether Muse Voice Transcribe will be built into its main consumer products.

  • Meta launches Muse Voice Transcribe, its first real-time audio transcription model
  • Distinguishes over 20 speakers and switches languages mid-sentence automatically
  • Arrives days after Google's rival Gemini 3.5 Transcribe model

AI Cricket Sport Technology

Read the full article at the source →

Originally published by Engadget as “Meta’s new AI transcription model can distinguish between multiple speakers and languages in real-time”.