OMX Helsinki — S&P 500 — DAX — NASDAQ 100 — STOXX 600 — EUR/USD — EUR/SEK — BTC/USD — ETH/USD — Euribor 3M — Euribor 12M —
AI Engineer

AI Engineer

Technology and AI 659,000 subscribers

Latency Is a Budget. Humanlike Is the Goal. — Jesse Hall, LiveKit

11.10.2026
Published: 11.10.2026 Category: Technology and AI

Episode details

Episode description

Your voice agent's dashboard says it's fast. Your callers say it feels robotic. They're both right. Jesse Hall, staff developer advocate at LiveKit, explains how to build voice agents that feel human. He shows why leaderboards can't pick your models (scores measure one model, but you ship a pipeline, and errors compound from speech-to-text to the LLM to text-to-speech) and why measured latency isn't what your callers hear. He covers buying time like a person does while tools run, plays a hotel-receptionist demo that handles interruptions and a caller changing his mind, explains LiveKit's audio-based turn detection (4.5% of users cut off vs 9.9% for the best alternative), and lays out the window that matters: a second and a half end to end still works, and around 600 milliseconds starts to feel human. He ends with LiveKit's open-source benchmark you can run against your own agent, a sneak peek of a voice-tuned Gemma model on LiveKit Inference, and a checklist. In this talk: • The leaderboard trap, and why pipelines fail differently than models • Measured vs perceived latency • Turn detection from audio, and buying time during tool calls • A latency budget, an open-source benchmark for your agent, and a checklist SPEAKER Jesse Hall, Staff Developer Advocate, LiveKit X: https://twitter.com/codeSTACKr LinkedIn: https://linkedin.com/in/codestackr LINKS LiveKit: https://livekit.io LiveKit Agents docs: https://docs.livekit.io/agents/ CHAPTERS 0:00 Intro 0:42 You've talked to one of these 1:22 Latency is a budget 2:02 Picking models 2:42 What LiveKit is 3:26 Realtime models vs pipelines 4:16 How pipeline errors compound 4:51 The leaderboard trap 6:06 Measured vs perceived latency 7:35 Buying time like a human 8:21 Demo: hotel receptionist 10:26 Audio-based turn detection 11:41 The 600 ms target 13:54 Benchmark your own agent 15:09 A faster voice LLM 15:53 The checklist Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #VoiceAI #LiveKit #AIEngineer

Open the AI assistant chat. The chat loads only when you open it.