Something interesting is happening inside Claude’s voice interface, and it’s a signal that tells us a lot about where AI assistants are heading.
For the past few weeks, if you’ve been living where the devs are and poking around with Claude’s voice mode, you might have noticed a model selector sitting there. It looked like you could choose between Opus, Sonnet, and Haiku. But here’s the thing: it was basically window dressing. No matter which option you clicked, you were still talking to Haiku 4.5, the same model that’s been running the show for a while now.
Now, when you select Opus or Sonnet, you actually get those models powering your voice conversations. This hasn’t officially been announced yet; it’s still hidden behind what’s known as a feature flag, but the functionality is there. And if history tells us anything, when the switches start working behind the scenes, a public launch is usually days away, not weeks.
I find this fascinating because it reveals something about the different paths companies are taking with voice AI.
OpenAI is going all-in on speech-native models by building systems that think in voice from the ground up. Anthropic is doing something completely different. They’re taking their most powerful reasoning models and plugging them into a conventional voice stack. The actual speech synthesis appears to be handled by ElevenLabs in the background, while Claude does what it does best: thinking through complex problems and holding intelligent conversations.
It’s the difference between rebuilding the engine and upgrading the one you’ve got.
For anyone who’s tried to have a serious conversation with Haiku in voice mode, this is going to feel like a massive leap. Haiku is fine for simple back-and-forth, but it gets thin quickly when you need deeper reasoning or want to tackle something complex through speech. Routing Opus or Sonnet through voice opens up entirely new possibilities. That means longer reasoning chains, tool-heavy requests, nuanced conversations that actually go somewhere.
I’m particularly intrigued by how this holds up in practice. From what I’ve read, the early testing suggests the interruption handling works well. The model pauses when you start speaking and picks back up naturally. It doesn’t jump in during long silences, which is crucial for making voice feel like an actual conversation rather than a race against an impatient assistant.
What strikes me most is the strategic choice here. While everyone else is chasing the shiny object of speech-native models, Anthropic is betting that what people really want is their smartest AI accessible through voice, even if the underlying architecture is more traditional. It’s a pragmatic approach: use the best reasoning models you have and make them available however people want to interact with them.
This reminds me of the early days of mobile apps, when everyone debated whether you needed native apps or if web apps would be good enough. The answer, as it turned out, wasn’t one-size-fits-all. Different approaches worked for different use cases, and the companies that won were the ones that focused on what users actually needed rather than what was technically “pure.”
The model picker currently shows Opus, Sonnet, and Haiku, with a thinking level control next to it. Notably absent is Claude Fable, their Mythos-class model, which isn’t available for voice yet. That might come later, or it might be a signal that certain models are better suited for certain interfaces.
No official announcement has dropped yet, just what has been updated in their help section, but the fact that the functionality is live behind the scenes tells you everything you need to know. This is coming, and soon.
For those of us who’ve been waiting for voice interfaces to get genuinely useful for complex work, this feels like a meaningful step forward. Not because it’s the flashiest approach or the most technically innovative, but because it takes the intelligence that already exists and makes it accessible in a new way.
Sometimes the best innovation isn’t inventing something entirely new. It’s taking what works and putting it where people need it.
I’ll be watching to see how this plays out when it goes public. The real test will be whether people find value in having Opus-level reasoning available through voice, or whether the speech-native approach ends up being what the market demands. My guess? There’s room for both, and we’re about to find out what different use cases call for.
The voice AI race is heating up, and the strategies couldn’t be more different. That’s when things get interesting.