
Voice AI can pass a Turing test. For about a minute.That's a generated clip, though. Have a human actually talk back and the number collapses to six or seven seconds, roughly where generated voice sat three years ago.One reason, per Aoden Teo of Miso Labs: real conversation isn't turn-based. Around 20% of the time more than one person is speaking, and laughter drives a lot of that overlap, since you laugh at a joke while it's still being told. We also adjust our pacing toward whoever we're talking to without noticing we're doing it.Voice models struggle with all of this. Full-duplex voice, where a model listens and speaks at the same time, is still extremely early.So an agent can know your joke is funny and still have to wait until you've finished before it laughs, by which point the timing has killed it.Aoden describes a second consequence: agents get pushed toward almost "psychotically emotive" behavior. If they can only talk once you've stopped, they need some other way to show they were listening. You finish your sentence, and the thing goes "Hmm?" You've heard it.Underneath that sits an architecture problem. Voice models have to respond fast, which constrains how large they can be, and fast means something different here than it does in text. Working with an LLM like Claude, Aoden points out, you care how quickly it finishes your code, not how quickly it starts.Voice inverts that. Nobody needs 10 hours of audio generated in two seconds, because nobody can listen to 10 hours of audio in two seconds; what matters is reaction time. Most architectural decisions trade latency against throughput, and Aoden expects voice to keep moving away from LLM-style designs toward ones built around very low latency.Miso is already pushing on it. Miso-1 got 3,000 stars on GitHub and 5 million views on Twitter, and they record data in their own LA studio because the internet doesn't contain every kind of audio a voice model might need. Nobody has released a podcast of someone reading millions and millions of email addresses, and people still want voice models that can read email addresses aloud, so teams end up generating some very strange training data themselves.The clip isn't the hard part. The hard part starts when you talk back."The most emotive foundation models for voice"🎙️Aoden Teo, CEO & Co-Founder, Miso Labs on Fondo START 1:03 Miso-1: 3K+ GitHub stars + 5M X views1:59 Why emotiveness matters for games, UGC + interactive products3:06 Measuring progress in voice AI with longer Turing tests4:01 Why interactive conversation is harder than generating convincing clips5:08 Full-duplex voice, interruptions + why laughter matters6:04 Latency vs. throughput - and why voice differs from LLMs7:09 Miso's LA recording studio + the challenge of voice training data9:02 Talking teddy bears, UGC, anime + unexpected voice AI use cases10:19 From serious chess player to math obsession to building @MisoLabsAI12:11 The surprise YC interviewCheck out misolabs.ai
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.
START: Jonathan Li, Founder & CEO, Quippy "Helps build social skills through daily practice"

START: Eric Chernoff, CEO & Founder, Fancysauce.ai "AI Cost Management, extra Fancy: Track usage, monitor ROI, and optimize spend across every AI workflow"

START: Ryaan Aqid, Founder & CEO, Quirk “ Infrastructure for information asymmetry”

START: Ilya Valmianski, CEO & Co-Founder, Signals "The AI store associate that turns buyers into regulars"
Free AI-powered recaps of START and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.