When an AI voice responds faster than a human would, it doesn't feel impressive — it feels creepy.

The latest voice AI models respond in under 150ms, while the average human takes 200 to 300ms to react.

One startup figured out that this 50ms gap is exactly what breaks trust — and they're betting $13 million on building the opposite.

TL;DR
Superhuman speed Distrust (the uncanny effect) Smallest.ai raises $13M Hybrid: parallel processing + LLM offloading A voice agent that "hesitates" like a human

Everyone Assumes Faster AI Voices Are Better

People building conversational AI have followed one rule of thumb for years: the lower the latency, the more human it sounds. The reasoning goes back to how humans actually talk — the natural gap between turns in a conversation is 200 to 500ms, and anything outside that range makes the exchange feel "broken".

So the industry spent the last few years racing to shrink that number. And it worked — the newest voice models now respond in about 150ms under streaming conditions. Here's the problem: that's faster than the average human reaction time of 200 to 300ms.

They won the speed race. The calls just got weirder.

Except That Speed Was the Problem

Voximplant CEO Alexey Aylarov calls this the "uncanny effect." His explanation: people read conversational timing as a signal — whether the other person is actually listening, rushing, or safe to talk to. When an AI cuts in with flawless, mechanically clean sentences, it comes across as ignoring those social cues entirely.

Key Takeaway

What actually builds trust isn't how low you get the latency — it's how human the pause in between feels.

That's exactly the gap Smallest.ai is targeting with this Series A. Its new architecture, "Hydra," doesn't process listening, reasoning, acting, and responding as a sequence — it runs them all in parallel. When a complex question comes in, instead of forcing an instant answer, it hands the question off to a larger LLM and naturally fills the wait with something like "let me look that up".

This works because they know exactly where the bottleneck sits — LLM inference accounts for 40 to 60% of total delay in a voice pipeline. Instead of hiding that delay, they turned it into the kind of pause a human makes when actually thinking — an "um, hang on a sec." The goal of this round isn't shaving off milliseconds at all costs — it's replicating the way humans naturally hesitate.

150ms
Latest voice AI response time
200–300ms
Average human reaction time
$21M+
Smallest.ai total funding raised

So Why Is Everyone Buying This Instead of Building It?

Smallest.ai founder Sudarshan Kamath sums up the answer: "Being world-class at voice is, for most companies, a distraction that pulls resources away from their actual business". Unless you're literally running call-center SaaS, tuning turn-taking timing and managing STT/TTS models isn't your job.

The numbers back this up. Smallest.ai says its customers have cut support costs by up to 80% and boosted agent productivity tenfold. RingCentral, Truecaller, and others are already on the customer list. The company sells this entire layer as a product too — TTS (Waves) and voice agents (Atoms), available as an API starting at $0.01 per minute.

This outsourcing trend shows up in the market-wide numbers too. VC investment in voice AI grew sevenfold, from $315 million in 2022 to $2.1 billion in 2024. In January 2026 alone, ElevenLabs raised $500 million, Decagon raised $250 million, and Deepgram raised $130 million. Not everyone's sold on it, though — Anthropic's Mike Krieger has said he still trusts text interfaces more.

The same trend is playing out elsewhere too. Gartner projects that agentic AI will be embedded in more than 33% of enterprise software by 2028, and large companies like Hyundai and Kakao are already integrating voice-based AI agents into their own services. It really comes down to one choice — build this layer yourself, or hand it to a specialist.

Line up three companies side by side, and the strategic differences become obvious.

ServicePositioningLatency Benchmark
Smallest.aiSmall specialized models + LLM hybridGenerates 10 seconds of TTS audio in 100ms
CartesiaSSM-based, supports on-premise deploymentSonic Turbo 40ms
ElevenLabsLeader in quality and variety, built for content creationFlash V2 75ms

Before You Plug In a Voice AI Agent, Check This

You don't need to build this kind of architecture yourself. But when you're picking a vendor, make sure you check these five things.

  1. Pin the latency requirement to an actual number
    Don't settle for "under 1 second" — ask for "300 to 500ms, with streaming support." A lot of adoption guides still cite "under 1 second" as the bar, which is way looser than the current industry standard.
  2. Interrupt the demo on purpose
    Test whether it yields naturally and picks back up, or just talks over you like a machine.
  3. Check the hand-off flow for complex questions
    When it delegates to an LLM, check whether the filler line sounds natural or whether you're just left with an awkward silence.
  4. Check on-premise and data-masking options
    Call recording and personal data handling vary a lot between vendors. Confirm upfront that it fits your deployment environment.
  5. Pilot with one small scenario first
    Don't hand over your entire support flow at once. Start with one high-repetition, low-risk scenario, watch the metrics, then expand.

Deep Dive Resources

Smallest.ai Official Site Check out the Waves (TTS) and Atoms (voice agent) product lines with live demos smallest.ai

The 300ms Rule (AssemblyAI) Breaks down why 300ms is the line voice AI can't cross, walking through the full 6-stage pipeline assemblyai.com

The Uncanny Effect in Modern Voice AI Voximplant's CEO breaks down why an AI voice that's too fast ends up feeling off-putting theaiinsider.tech

Cartesia vs. ElevenLabs A direct comparison of how SSM-based and transformer-based architectures stack up on latency cartesia.ai

Voice AI in 2026 (AssemblyAI) A data-driven series tracking voice AI investment and adoption trends — start here for the full market picture assemblyai.com

AI Voice Agent Adoption Guide (Humelo) Covers the difference between callbots and voice agents, plus the actual adoption checklist Korean companies use humelo.com