Support Tech / Meta Launches 'Muse Voice Transcribe' with Indian Language Support

Key Points
Meta unveils Muse Voice Transcribe, a breakthrough real‑time audio model with adaptive delay, multilingual support including five Indian languages, and advanced streaming transcription, ranking first on global speech‑to‑text benchmarks.
New Delhi, Sep 2: Meta has stepped into real‑time audio perception with the launch of Muse Voice Transcribe, a breakthrough model built by its Superintelligence Labs.
Designed to transform speech‑to‑text technology, the system offers advanced streaming transcription, seamless code‑switching, and native support for five major Indian languages.
It can separate more than 20 voices in hour‑long recordings and perform diarisation — all within a single model, eliminating the need for post‑processing.
Also read:
📱 Get Argus News App
✨Trained across more than 70 languages, with 25 validated at launch, Muse Voice Transcribe has already achieved top ranking. “It ranks first on the Artificial Analysis streaming speech‑to‑text leaderboard as of September 1, 2026,” Meta said in a statement.
The model is available via the Meta Model API and is already powering dictation in Meta AI for Mac and Muse Code. It offers real‑time automatic speech recognition, diarisation with over 20 speakers, and endpointing. Accuracy is enhanced through language, keyword, and context biasing.
A key innovation is its “adaptive delay” mechanism. “The longer the model waits to predict, the more accurate the transcript, but the higher the latency. Muse Voice Transcribe has ‘adaptive delay,’ dynamically changing delay for each word based on difficulty,” the company explained.
This feature is enabled through reinforcement learning, combining word error rate (WER) reward with delay reward multiplicatively.
Meta described Muse Voice Transcribe as an autoregressive multimodal model from the Muse Spark family. Audio is processed in 80‑millisecond chunks, each converted into a soft token.
At every step, the model decides whether to continue listening or emit a text token. With adaptive delay, it achieves the Pareto front on speed‑accuracy trade‑off, measured by time to final transcription.
(IANS)
Related Topics
Explore more stories