Meta has expanded its push into voice-based artificial intelligence with Muse Voice Transcribe, a real-time speech-to-text model designed to transcribe conversations while people are speaking and handle multilingual conversations within a single system.
Developed by Meta Superintelligence Labs (MSL), the model supports five major Indian languages — Hindi, Tamil, Telugu, Kannada and Malayalam — alongside broader multilingual capabilities.
Meta says Muse Voice Transcribe was trained across more than 70 languages, with 25 languages extensively validated for its initial release.
Beyond converting speech into text, the model combines real-time automatic speech recognition, speaker identification and endpoint detection while also handling code-switching — an important capability in multilingual markets such as India, where speakers frequently move between English and regional languages during the same conversation.
What Is Muse Voice Transcribe?
Muse Voice Transcribe is Meta Superintelligence Labs' first real-time audio perception model.
Instead of waiting for an entire audio recording to finish before generating a transcript, the system produces text continuously as speech arrives.
The model combines several functions that can otherwise require separate processing stages.
It provides:
Real-time streaming automatic speech recognition
Speaker diarisation for more than 20 speakers
Endpointing to recognise speech boundaries
Multilingual transcription
Seamless code-switching
Language, keyword and context biasing
Support for recordings exceeding one hour
Meta says these capabilities operate natively within the model without requiring a separate post-processing step.
Five Indian Languages Supported
The Indian-language support is one of the most significant elements of the launch for Meta's users and developers in the country.
Muse Voice Transcribe supports:
Hindi, Tamil, Telugu, Kannada and Malayalam.
The model was trained across more than 70 languages overall, although Meta says 25 have been extensively verified for the initial release.
That distinction is important: training coverage does not necessarily mean every language has received the same degree of validation.
For India, however, the five supported languages could make the technology useful for applications ranging from voice assistants and meeting transcription to customer-service systems, dictation and other speech-driven software.
Code-Switching Could Be Particularly Useful in India
Muse Voice Transcribe's ability to recognise code-switching may prove more important than simply increasing the number of supported languages.
Code-switching occurs when speakers move between languages during a conversation — sometimes even within the same sentence.
An Indian user, for example, might begin a sentence in English, move into Hindi and then return to English without consciously treating them as separate conversations.
Traditional speech-recognition systems can struggle with such changes or require explicit language selection.
Meta says Muse Voice Transcribe can process multilingual input and recognise these switches without requiring separate models.
If that performance translates reliably to everyday conversations, the capability could make real-time transcription more practical in multilingual markets.
More Than 20 Speakers in Hour-Long Recordings
Meta is also targeting more complicated audio environments rather than limiting the model to one-to-one conversations.
Muse Voice Transcribe can perform speaker diarisation involving more than 20 speakers and process audio lasting longer than one hour, according to Meta.
Speaker diarisation means identifying which portions of a conversation belong to different speakers.
That capability could be useful for meetings, conferences, interviews, panel discussions, lectures and other recordings where a transcript needs to distinguish between multiple participants.
Meta says these features are handled natively by the model rather than relying on an additional post-processing system after transcription.
Adaptive Delay Balances Speed and Accuracy
Real-time transcription creates a fundamental technical challenge.
A speech-recognition model can wait longer before deciding what someone said, potentially improving accuracy — but doing so increases the delay before text appears.
Producing text too quickly can reduce latency but potentially increase errors.
Meta says Muse Voice Transcribe addresses that trade-off through a feature called adaptive delay.
Instead of using the same delay for every word, the model can dynamically determine how long it needs to listen depending on the difficulty of recognising a particular word.
Relatively straightforward speech can therefore be processed quickly, while more difficult portions can receive additional processing time.
Meta says reinforcement learning is used to optimise this balance between transcription error and delay.
Audio Processed in 80-Millisecond Chunks
Meta has also provided technical details about how the model processes incoming speech.
Muse Voice Transcribe processes audio in 80-millisecond chunks, or 12.5 Hz.
Each audio chunk is transformed into a soft token.
At every stage, the model effectively chooses between continuing to listen to additional audio or producing a text token.
This architecture is designed to allow transcription to appear progressively while retaining enough context to recognise more difficult speech.
Meta describes Muse Voice Transcribe as an autoregressive multimodal model from the Muse Spark family.
Meta Claims No. 1 Streaming Speech-to-Text Ranking
Meta says Muse Voice Transcribe ranked first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026.
The timing and attribution are important.
This is a performance claim reported by Meta based on the leaderboard's results at the time of launch; it should not be interpreted as proof that the system will outperform every competing transcription service under every real-world condition.
Performance can vary considerably depending on accents, microphones, background noise, language combinations, speaker overlap and the type of audio being processed.
Real-world performance across India's wide range of accents and multilingual speech patterns will therefore be an important measure of the technology beyond benchmark results.
Already Available in Meta AI for Mac and Muse Code
Muse Voice Transcribe is not merely a research demonstration.
Meta says the model is already powering voice dictation in Meta AI for Mac and Muse Code.
Developers can also access the technology through the Meta Model API, potentially allowing third-party applications to integrate its speech-recognition capabilities.
Meta has priced API access at $3 per 1,000 audio minutes, equivalent to approximately $0.18 per hour of audio.
That gives developers a direct route to build applications around the technology without developing their own speech-recognition infrastructure.
Why the Launch Matters for India's AI Market
India presents an unusually demanding environment for speech-based AI.
The challenge is not simply supporting a large number of languages. Real conversations can involve regional accents, English vocabulary mixed into Indian languages, multiple speakers and rapid switching between languages.
That makes code-switching and multilingual recognition particularly relevant.
Muse Voice Transcribe's five Indian languages do not cover India's entire linguistic landscape, and Meta has not claimed that they do.
However, combining regional-language support with real-time transcription and code-switching represents a more sophisticated approach than treating each language as an isolated speech-recognition task.
For businesses, the technology could eventually support multilingual meeting transcription, customer service, accessibility tools, voice assistants and productivity applications.
For consumers, the immediate application is more visible through dictation within Meta's own AI products.
Meta's Broader Push Into Indian-Language AI
The launch also fits into Meta's wider effort to make AI-powered communication work across Indian languages.
Meta has previously expanded AI-powered translation for Reels on Instagram and Facebook to additional Indian languages.
Muse Voice Transcribe tackles a different part of the language problem: understanding spoken audio and turning it into text quickly enough for interactive applications.
The distinction matters because speech recognition can serve as an underlying layer for several types of AI products.
Once an AI system can reliably understand live multilingual speech, that capability can potentially feed into assistants, translation systems, coding tools, meeting software and other applications.
What Comes Next
The biggest test for Muse Voice Transcribe will be performance outside controlled benchmarks.
Meta's claimed leaderboard position and technical capabilities make the launch notable, but Indian-language speech recognition involves considerable real-world complexity.
Background noise, overlapping speakers, regional accents and conversations that continuously switch between languages will test whether the model's benchmark performance translates into everyday reliability.
For now, Meta has delivered an important addition to its AI stack: a real-time speech model that combines transcription, speaker separation and multilingual understanding while giving five Indian languages a place in its initial validated language set.
For India's rapidly growing AI ecosystem, the most consequential feature may ultimately be not simply that Muse can understand several languages — but that it is designed to understand how multilingual people actually speak.






