The company’s first real-time audio model can handle 20+ speakers, switch between languages mid-sentence, and even catch that thing where you mix English and Spanish without breaking a sweat
Look, I’ve been covering AI for a minute now, and I’ve seen plenty of “breakthroughs” that turn out to be hype wrapped in a press release. But every once in a while, something drops that actually makes you sit up and pay attention.
Meta just dropped Muse Voice Transcribe, and honestly? It’s kind of a big deal.
This isn’t your grandma’s speech-to-text. We’re talking about a real-time audio model that can handle dictation and transcription for more than 20 speakers at once — and seamlessly switch between multiple languages in the same conversation. Yeah, you read that right. Twenty-plus people, multiple languages, all happening in real time.
Mark Zuckerberg himself — who, by the way, just came back to X after three years of radio silence — posted a demo video showing off what this thing can do. And the clip is pretty wild. The transcription automatically distinguishes between different speakers, switches between languages on the fly, and even picks up on “code-switching” — you know, that thing where bilingual speakers mix two languages in the same sentence. It transcribes that too. No glitches. No awkward pauses. Just seamless, real-time understanding.
This is Meta’s first real-time audio perception model from its Superintelligence Lab (MSI), and it’s rolling out today. But before we get into the weeds, let’s talk about why this actually matters — and why it’s got the AI world buzzing.

The Tech Under the Hood: Why This Isn’t Just Another Transcription Tool
So what makes Muse Voice Transcribe different from the countless other speech-to-text models out there?
For starters, it’s state-of-the-art in streaming speech-to-text — handling speaker diarization (that’s the fancy term for “figuring out who said what”) and endpointing natively in a single model. Most transcription tools use separate systems for speaker identification and transcription, which adds latency and creates room for error. Meta’s doing it all in one shot.
But here’s the really clever part: the model decides when to listen. Zuckerberg explained it this way: “It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy”. Think about that for a second. Instead of processing audio at a fixed rate — which is how most transcription works — this thing is dynamically adjusting its attention based on how difficult the incoming speech is. Hard word? It pauses, takes a beat, processes more context. Easy word? It moves on immediately.
That’s not just a marginal improvement. That’s a fundamentally smarter way to do real-time transcription.
And the model holds up on messy, real-world audio too. We’re not talking about pristine studio recordings with professional microphones. We’re talking about background noise, overlapping conversations, people talking over each other — the kind of chaotic audio that makes traditional transcription models throw up their hands and give you gibberish.
Meta trained Muse Voice Transcribe across more than 70 languages, with 25 validated and ready to go at launch. That’s a massive linguistic footprint, and it means the model can handle conversations that span multiple languages without missing a beat.
The Demo That Got Everyone Talking
Zuckerberg’s demo video — which, if we’re being honest, is probably the most interesting thing he’s posted since his return to X — shows the model in action. And it’s genuinely impressive.
You’ve got multiple speakers, different voices, different languages, and the transcription is keeping up in real time. It’s not just transcribing words; it’s attributing them to the right person, flagging who’s speaking when, and handling language switches so smoothly you barely notice it’s happening.
The code-switching capability is particularly noteworthy. If you’ve ever spent time in bilingual communities — whether it’s Spanglish in Miami, Taglish in Manila, or Hinglish in Mumbai — you know that code-switching is natural, fluid, and almost impossible for traditional transcription tools to handle. Most models get confused when you switch languages mid-sentence. They either drop words, misinterpret context, or just give up entirely.
Muse Voice Transcribe handles it like it’s no big deal. That’s a genuine technical achievement.
The Race Is On: Meta vs. Google
Here’s where things get interesting. Meta’s release comes less than a week after Google unveiled Gemini 3.5 Transcribe, its own audio model with similar capabilities .
Google’s offering is pretty impressive too. The company says Gemini 3.5 Transcribe can automatically detect more than 85 languages, adapt unstructured speech into formatted text, remove filler words, and even learn custom vocabulary and unique spellings . It can attribute speech to up to three speakers with word-level timestamps based on pre-recorded audio. That’s solid for podcast transcription and similar use cases.
But there’s a key difference in how these two companies are approaching the market.
Google is baking its audio model directly into Android and, eventually, Chrome. You’ll be able to use speech-to-text in any web field in Chrome — dictate replies, posts, prompts, whatever. It’s already powering the Rambler feature on Android devices like the Pixel 11-series phones, as well as the Gemini app on macOS. Google is embedding this deeply into its ecosystem, making it a core part of how users interact with its products.
Meta, on the other hand, hasn’t said whether it plans to integrate Muse Voice Transcribe into its flagship services like WhatsApp, Messenger, or Instagram. For now, the model is available through the Meta AI Mac app — which, because it can power voice-enabled features on other apps, will let Muse Voice Transcribe handle dictation across other services. But that’s a very different play than Google’s ecosystem-first approach.
It’s worth noting that this isn’t just about transcription. Both companies are positioning audio AI as a foundational capability for the next generation of human-computer interaction. Google explicitly frames Gemini 3.5 Transcribe as understanding “your natural intent and speaking style”. Meta’s model is part of a broader push from MSI, which has been churning out new AI models at a remarkable pace.
The Bigger Picture: Meta’s AI Factory Is Running Hot
Muse Voice Transcribe is the latest release from Meta Superintelligence Lab, and it’s arriving amid a flurry of activity.
In just the last few weeks, Meta has also introduced Muse Code — its take on a coding agent that competes directly with Anthropic’s Claude Code and OpenAI’s Codex. That tool is powered by Muse Spark 1.2 and can handle software engineering tasks like writing code, planning changes, and validating results. It can even manage multiple sub-agents and delegate tasks. And Meta is pricing it aggressively — $1.25 per million input tokens and $4.25 per million output tokens, with a “contributor tier” that drops to just $0.10 per million input tokens for users who provide feedback. For context, Anthropic’s Sonnet 5 costs $3 per million input tokens and $15 per million output tokens. That’s a massive price difference.
The company also released Muse Glimmer, a slimmed-down “open source” model that’s light enough to run on a single computer. Based on Meta’s Spark 1.2 closed model, Glimmer requires just a single GPU for agent-oriented tasks like scheduling and file management. Meta is making the weights available for free on Hugging Face, and users can run the model on their own PCs. Zuckerberg framed it as part of a broader philosophy: “Rather than centralizing superintelligence, we should distribute it widely and give every person the ability to direct it”.
Put all of this together, and you start to see a pattern. Meta is building out an entire AI ecosystem — coding agents, lightweight local models, and now real-time audio perception. The company is competing on multiple fronts simultaneously, and it’s doing so with aggressive pricing and a willingness to open-source key components.
What This Means for Developers and Users
For developers, Muse Voice Transcribe is available now through Muse Code and Meta’s Model API, priced at $3 for 1,000 audio minutes. That’s not cheap, but it’s competitive with similar offerings — especially given the model’s capabilities.
There’s also a demo version available on Meta’s research blog, so you can kick the tires before committing. And because the model is already powering dictation in the Meta desktop app and Muse Code, developers can start building with it immediately.
For everyday users, the impact is more immediate than you might think. The Meta AI Mac app can power voice-enabled features on other apps, which means Muse Voice Transcribe could soon be handling dictation across a wide range of services. Imagine being able to dictate emails, documents, and messages with a tool that actually understands who’s speaking, handles multiple languages, and adapts to difficult words in real time. That’s not a futuristic vision — that’s available today.
The question, of course, is whether Meta will eventually integrate this into its core products. WhatsApp, Messenger, and Instagram have billions of users combined. If Muse Voice Transcribe becomes the default transcription engine for voice messages and calls on those platforms, it could fundamentally change how people communicate. But for now, Meta is staying quiet on that front.
The Privacy Question Nobody’s Talking About — Yet
Let’s be real for a second. A model that can transcribe conversations with 20+ speakers in real time, across multiple languages, with speaker identification built in — that’s powerful. And with great power comes great responsibility, or however the saying goes.
Meta hasn’t said much about privacy implications yet. The model is available through the Meta AI Mac app, which means audio data is presumably being processed in the cloud. That raises questions about data retention, user consent, and whether Meta plans to use transcribed conversations for training future models.
Google, for its part, is framing Gemini 3.5 Transcribe as a tool that “understands your natural intent and speaking style”. That’s a nice way of saying the model is learning from your speech patterns. Whether that data is anonymized, how long it’s stored, and who has access to it — these are all open questions.
The industry as a whole is still figuring out the rules of engagement for audio AI. Unlike text-based interactions, which leave a clear digital trail, voice conversations are more intimate and potentially more sensitive. A model that can identify individual speakers and transcribe their words in real time is essentially creating a permanent record of spoken conversations. That’s useful, but it’s also concerning.
Meta and Google both have checkered histories when it comes to user privacy. Neither company has earned the benefit of the doubt. So while the technology is impressive, it’s worth keeping a skeptical eye on how it’s deployed and what safeguards are in place.
The Competitive Landscape: Who’s Winning the Audio AI Race?
It’s too early to declare a winner, but we can already see the contours of the competition.
Google has the ecosystem advantage. Gemini 3.5 Transcribe is baked into Android, coming to Chrome, and integrated with the Gemini app on macOS. That means Google’s model will be the default choice for millions of users who never have to think about which transcription tool they’re using — it’s just there, working in the background.
Meta has the research advantage and the aggressive pricing. The company is releasing models at a remarkable pace — coding agents, local models, audio perception — and it’s pricing them to compete. The $3 per 1,000 audio minutes for Muse Voice Transcribe is reasonable, and the broader trend of Meta undercutting competitors on price is hard to ignore.
Then there’s the open-source angle. Meta’s release of Muse Glimmer as a free, downloadable model that runs on a single computer is a direct challenge to the closed, centralized approach of OpenAI and Anthropic. If developers can run powerful AI models locally, without sending data to the cloud, that changes the calculus entirely.
The Chinese AI companies — DeepSeek and others — are also in the mix, offering open licensing and local deployment options at even lower prices. The market is fragmenting, and that’s probably a good thing for users. Competition drives innovation and keeps prices in check.
What’s Next for Muse Voice Transcribe?
For now, the model is available through the Meta AI Mac app, Muse Code, and Meta’s Model API. Developers can start building with it immediately. Users can experience it through the desktop app.
But the real story is what comes next.
Will Meta integrate this into WhatsApp, Messenger, and Instagram? That would be a game-changer for how billions of people communicate. Will the model improve over time? Meta says it’s trained across 70+ languages, with 25 validated at launch — that number is likely to grow. Will we see a mobile version? Nothing announced yet, but it’s hard to imagine Meta keeping this desktop-only for long.
And what about the broader implications? A model that can handle 20+ speakers in real time, across multiple languages, with adaptive delay and native speaker diarization — that’s not just a transcription tool. That’s a foundational capability for ambient computing, for AI that lives in the background and understands what’s happening around you.
We’re not quite there yet. But Muse Voice Transcribe is a significant step in that direction.
The Bottom Line
Meta just released its first real-time audio model, and it’s genuinely impressive. Muse Voice Transcribe can handle more than 20 speakers, seamlessly switch between languages, pick up on code-switching, and adapt its processing speed based on word difficulty. It’s state-of-the-art in streaming speech-to-text, and it’s available now through the Meta AI Mac app and Meta’s API.
The timing is interesting. Meta’s release comes less than a week after Google unveiled Gemini 3.5 Transcribe, setting up a direct competition between two tech giants . Google is embedding its model deeply into its ecosystem; Meta is making its model available through developer tools and desktop apps, with no clear integration plans for its flagship services yet.
Both approaches have merit. Google’s ecosystem play means more users will encounter the technology without having to seek it out. Meta’s developer-first approach means more experimentation and innovation from third-party builders.
Either way, audio AI is having a moment. And Muse Voice Transcribe is proof that the technology is advancing faster than most people realize. It’s not perfect — privacy questions remain unanswered, and we’re still waiting to see how Meta plans to scale this — but it’s a genuine step forward.
For now, if you’ve got a Mac and you’re curious about what real-time, multi-speaker, multi-language transcription looks like, go check out the Meta AI app. The demo on Meta’s research blog is worth a look too.
Just remember: the model is listening. And it actually knows who’s talking.
