<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[RealTime AI - Weekly Updates]]></title><description><![CDATA[RealTime AI - Weekly Updates about Voice Agents, Speech Technology and Conversational Models ]]></description><link>https://livetok.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png</url><title>RealTime AI - Weekly Updates</title><link>https://livetok.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 02:31:50 GMT</lastBuildDate><atom:link href="/__u/livetok.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Gustavo]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[livetok@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[livetok@substack.com]]></itunes:email><itunes:name><![CDATA[Gustavo]]></itunes:name></itunes:owner><itunes:author><![CDATA[Gustavo]]></itunes:author><googleplay:owner><![CDATA[livetok@substack.com]]></googleplay:owner><googleplay:email><![CDATA[livetok@substack.com]]></googleplay:email><googleplay:author><![CDATA[Gustavo]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Weekly Updates - Aug 31st 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-aug-31st-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-aug-31st-2026</guid><pubDate>Mon, 31 Aug 2026 14:14:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>&#128478;&#65039; Market and Product News</h2><ul><li><p><a href="https://x.com/mati/status/2092254456549167152">Stripe becomes an ElevenLabs customer, running support through ElevenAgents</a><br>ElevenLabs co-founder Mati Staniszewski says Stripe is now using ElevenAgents to interact with customers, on top of ElevenLabs already running its own billing through Stripe.</p></li><li><p><a href="https://techcrunch.com/2026/08/25/indias-ringg-gets-backing-from-peak-xv-as-it-pushes-voice-ai-past-the-phone-call/">India&#8217;s Ringg raises $10M Series A extension from Peak XV to push voice AI past the phone call</a><br>The Bengaluru startup now processes 20M call attempts a month for customers like Flipkart and CRED, and is expanding its voice agents into WhatsApp and browser workflows.</p></li></ul><h2>&#129520; Platform News</h2><ul><li><p><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/">Google ships Gemini 3.5 Transcribe, 2.6% WER across 85+ languages</a><br>The model streams in real time via a Live API, already wired into LiveKit, Pipecat, and Agora, and adds speaker attribution for up to three speakers.</p></li><li><p><a href="https://www.daily.co/blog/announcing-pipecat-phonellm-alpha-1/">Daily open-sources PhoneLLM, a voice-agent LLM at 1/18th the cost of GPT-5.6 Terra</a><br>The Nemotron-3-Nano fine-tune matches GPT-5.6 Terra on typical voice agent tasks at roughly 1/3 the latency, alongside a new PhoneBench v1 benchmark for phone-agent LLMs.</p></li><li><p><a href="https://turnbench.sesame.com/">Sesame launches TurnBench, an open benchmark for conversational turn-taking</a><br>The public benchmark pairs a 30-hour hand-labeled conversation corpus with a leaderboard, interactive viewer, and self-serve scoring for end-of-turn and interruption detection.</p></li><li><p><a href="https://www.tavus.io/blog/sparrow-2">Tavus ships Sparrow-2, a real-time conversational understanding model</a><br>Sparrow-2 models semantics, prosody, and background noise to decide when to listen, wait, or speak &#8212; built for noisy real-world settings like caf&#233;s, airports, and retail floors.</p></li><li><p><a href="https://www.cartesia.ai/blog/sonic-3.6">Cartesia&#8217;s Sonic-3.6 tops both Artificial Analysis speech arenas</a><br>The rebuilt TTS model adds Odia and Urdu for 44 total languages and wins blind naturalness tests against Sonic-3.5 up to 93% of the time.</p></li><li><p><a href="https://livekit.com/blog/introducing-per-second-metering-for-livekit">LiveKit switches to per-second billing</a><br>Agent sessions, recording, WebRTC, and SIP now bill in 10-second increments instead of full minutes, with no code changes required.</p></li><li><p><a href="https://www.retellai.com/blog/retell-workflows">Retell AI launches Workflows for native CRM and helpdesk automation</a><br>Retell Workflows lets voice agents update CRM and helpdesk systems mid-call instead of relying on post-call webhooks.</p></li><li><p><a href="https://elevenlabs.io/blog/elevenlabs-cli-v1">ElevenLabs CLI v1 brings agents-as-code to the terminal</a><br>Every ElevenLabs API operation is now a subcommand with JSON/table/YAML output, a --dry-run preview mode, and config files for editing ElevenAgents like code.</p></li><li><p><a href="https://www.assemblyai.com/blog/qwen-4b-llm-gateway">AssemblyAI adds Qwen3.5 4B to its LLM Gateway for voice rewrite tasks</a><br>The hosted model averages 612ms on dictation cleanup and live formatting &#8212; 1.9x faster than GPT-4.1 at 94% lower cost.</p></li></ul><h2>&#128214; Reading</h2><ul><li><p><a href="https://www.nojitter.com/ai-automation/real-time-ai-coaching-is-becoming-continuous-surveillance">Real-time AI coaching is becoming continuous surveillance</a><br>No Jitter argues live AI coaching tools for contact-center agents risk sliding from support into constant performance surveillance without clear guardrails.</p></li><li><p><a href="https://vapi.ai/blog/build-vs-buy">Vapi: the build-vs-buy voice agent decision is really about orchestration</a><br>Argues teams underestimate the orchestration layer&#8217;s complexity and are usually better off buying a platform when voice is a feature, not the core product.</p></li><li><p><a href="https://bloggeek.me/gpt-live-full-duplex-voice-ai/">OpenAI just changed how Voice AI is built</a><br>Tsahi Levent-Levi on GPT-Live-1&#8217;s shift to full-duplex streaming: WebRTC call setup drops from 6 round trips to 2, and turn detection gets replaced by continuous duplex audio.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p><span>LiveKit Agents: </span><a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.7.1">v1.7.1</a><span>. </span>Gemini 3.5 Transcribe, Palabra STT and TTS, Sarvam improvements and many fixes.</p></li><li><p><span>Pipecat: </span><a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.8.0">v1.8.0</a><span>. </span>This is a massive release with improvements across reliability, startup performance, turn management, evals, function calling, services, transports, and developer tooling. There is a lot in this one!</p></li></ul><ul><li><p>TEN Framework: No releases.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RealTime AI - Weekly Updates! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jul 27th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jul-27th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jul-27th-2026</guid><pubDate>Tue, 28 Jul 2026 21:56:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>&#128478;&#65039; Market and Product News</h2><ul><li><p><a href="https://techcrunch.com/2026/07/24/openais-new-voice-mode-makes-it-to-the-chatgpt-desktop-app/">ChatGPT Voice comes to desktop, can drive Codex hands-free</a>. OpenAI&#8217;s GPT-Live-powered voice mode dictates in any window and can tell Codex to create branches, open pull requests, and debug &#8212; live on macOS and Windows.</p></li><li><p><a href="https://openai.com/index/introducing-openai-presence">OpenAI launches Presence, an enterprise voice/chat agent platform</a>. Presence&#8217;s live phone line (1-888-GPT-0090) resolves 75% of inbound issues without human escalation, with early customers BBVA Mexico, SoftBank, and IAG.</p></li><li><p><a href="https://www.pymnts.com/news/investment-tracker/2026/startup-telli-raises-15-million-to-replace-call-centers-with-ai-agents/">Telli raises $15M seed to replace call centers with AI agents</a>. Redalpine led the round, with Y Combinator and Cherry Ventures joining, backing Telli&#8217;s AI agent &#8220;Charlie&#8221; that answers calls and books appointments for clients like Sky and Vaillant.</p></li><li><p><a href="https://www.zoom.com/en/blog/voice-translation-zoom/">Zoom launches live voice translation in meetings</a>. Zoom&#8217;s Voice Translator chains ASR, machine translation, and TTS to translate speech live across English, Chinese, Spanish, French, and Japanese, with Arabic coming later in 2026.</p></li><li><p><a href="https://hothardware.com/news/apple-genius-bar-ai-recording-tool">Apple pilots AI transcription at the Genius Bar</a>. Live Notes records in-store appointments only with both parties&#8217; consent, auto-generating an editable summary walled off from managers during the pilot.</p></li><li><p><a href="https://www.smartcitiesworld.net/news/ukraine-adds-voice-ai-to-national-government-assistant-13027">Ukraine adds ElevenLabs voice AI to its Diia government app</a>. Diia.AI now lets citizens access 170+ government services through spoken conversation, built with GovTech firm Kitsoft, with seamless mid-interaction switching between voice and text.</p></li><li><p><a href="https://www.tipranks.com/news/private-companies/elevenlabs-voice-ai-gains-traction-through-sevenrooms-restaurant-deployment">SevenRooms&#8217; ElevenLabs voice agents pass 400,000 restaurant calls</a>. SevenRooms, a DoorDash company, has handled 400,000+ calls with ElevenAgents across 600+ restaurant venues, helping them capture up to 25% more reservations.</p></li><li><p><a href="https://www.nojitter.com/contact-centers/ringcentral-strong-2q-2026-powered-by-customer-growth-across-several-products">RingCentral&#8217;s AI receptionist customers quadruple year-over-year</a>. RingCentral&#8217;s AI Receptionist base grew to 16,400 customers and AI Conversation Expert grew 70% to 6,300 customers, on total Q2 revenue of $657M.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/2-billion-calls-parlance-voice-ai-nears-healthcare-milestone-302832861.html">Parlance nears 2 billion healthcare voice AI calls</a>. The platform has answered 1.9 billion+ patient calls at 6+ per second, sustaining 87% caller engagement across 400+ health systems in the US, Canada, and UK.</p></li></ul><h2>&#129520; Platform News</h2><ul><li><p><a href="https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/">Alibaba&#8217;s Qwen-Audio-3.0-TTS tops the speech arena at 1,236 Elo</a><br>The Plus tier hit #1 on Artificial Analysis&#8217;s TTS arena, while the real-time Flash tier delivers 300ms first-packet latency across 16 languages and 20 Chinese dialects.</p></li></ul><ul><li><p><a href="https://www.speechmatics.com/company/articles-and-news/speaker-lock-fixing-voice-ai-for-the-real-world">Speechmatics adds speaker-aware &#8220;Speaker Focus&#8221; to LiveKit agents</a><br>The diarization feature dynamically locks onto a target speaker to filter background noise and crosstalk in scenarios like drive-thrus and noisy support calls.</p></li><li><p><a href="https://langchain.com/blog/trace-voice-agents-in-langsmith">LangSmith adds tracing for four voice agent frameworks</a><br>LangSmith now traces Pipecat, LiveKit, OpenAI Realtime, and Gemini Live agents with full conversation audio, STT/TTS latency, interruption detection, and tool calls.</p></li><li><p><a href="https://vapi.ai/blog/model-intelligence">Vapi launches Model Intelligence presets and performance metrics</a><br>Four curated presets &#8212; Balanced, High Intelligence, Ultra Fast, Cost Saver &#8212; bundle transcriber, model, and voice, backed by live latency/cost/quality data from 1B+ calls on the platform.</p></li><li><p>Updates for different models in the Coval&#8217;s benchmarks. <a href="https://www.assemblyai.com/blog/universal-3-5-pro-realtime-human-parity-zone">AssemblyAI&#8217;s Universal-3.5 Pro Realtime STT hits Coval&#8217;s &#8220;Human Parity Zone&#8221;</a>, <a href="https://www.coval.ai/blog/tts-latency-benchmark-2026">Palabra TTS v1 becomes the fastests model</a> and <a href="https://x.com/inworld_ai/status/2079969059572199812">Inworld claims the top spot on a real-world STT benchmark</a>.</p></li></ul><ul><li><p><a href="https://getbluejay.ai/">Bluejay opens self-serve signup for its voice agent QA platform</a><br>The YC-backed platform, already used by Google and Zocdoc, now lets anyone test and improve voice and chat AI agents with no credit card required.</p></li></ul><h2>&#128214; Reading</h2><ul><li><p><a href="https://www.coval.ai/blog/tts-latency-benchmark-2026">TTS Latency 2026: The Production Speed Metric for Voice AI</a> (Coval). Human conversation requires a response time under 400 milliseconds to feel natural, minimizing text-to-speech (TTS) latency is critical for voice agents, a benchmark currently led by Palabra TTS v1 with an industry-low 103-millisecond response time.</p></li><li><p><a href="https://www.forbes.com/sites/frankracioppi/2026/07/22/is-ai-taking-over-audio---radio-podcasting-audiobooks/">Forbes asks whether AI is taking over radio, podcasts, and audiobooks</a>. Podcast Index data shows 35-45.7% of newly uploaded podcasts appear AI-generated, and an Edison study found AI-narration trials nearly doubled listener interest, from 31% to 65%.</p></li><li><p><a href="https://www.cxtoday.com/contact-center/why-cx-leaders-should-stop-measuring-voice-ai-by-deflection-alone-parloa-cs-0228/">Stop grading voice AI on deflection alone</a> (Parloa) Deflection-only metrics reward hanging up on customers &#8212; Parloa argues teams should track resolution rate, customer effort, sentiment shift, and escalation quality instead.</p></li><li><p><a href="https://vapi.ai/blog/how-to-reduce-wait-times-with-ai">How to reduce call center wait times with AI</a> (VAPI)  Cites Freshworks data showing AI agents deflect 45%+ of incoming queries, and an NBER study finding generative-AI assistance raised agent throughput 13.8%.</p></li></ul><h2><strong>&#128230; Releases</strong></h2><ul><li><p><span>LiveKit Agents: </span><a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.7">v1.6.7</a><span>. Spatius avatars, Voice Isolation Krisp mode, telemetry updates and many improvements and fixes.</span></p></li><li><p><span>Pipecat: </span><a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.6.0">v1.6.0</a>. New Media over QUIC Transport.<span> Deepgram Flux TTS (early-access streaming TTS), plus Baseten and Crusoe Cloud LLM services for open-weights models, reasoning for OpenAI Responses better OTel Tracing and many small fixes and improvements.</span></p></li></ul><ul><li><p>TEN Framework: No releases.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RealTime AI - Weekly Updates! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jul 22nd 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jul-22nd-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jul-22nd-2026</guid><pubDate>Mon, 20 Jul 2026 19:42:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>&#128478;&#65039; Market and Product News</h2><ul><li><p><a href="https://techcrunch.com/2026/07/15/rime-picks-up-24m-series-a-to-help-enterprises-field-customer-calls/">Rime raises $24M to build enterprise speech-to-speech models</a>. M13-led Series A backs models already powering ~100M monthly calls for Mayo Clinic, Dialpad, and Asurion; new hire Rafael Valle (ex-Meta, NVIDIA) joins as chief scientist.</p></li><li><p><a href="https://www.pwc.com/us/en/about-us/newsroom/press-releases/pwc-openai-agentic-contact-service-solutions.html">PwC and OpenAI team up on agentic customer service</a>. PwC&#8217;s voice and digital agent stack, built on OpenAI&#8217;s multimodal APIs, gets a joint Center of Excellence to help enterprises deploy it across sales, marketing, and service.</p></li><li><p><a href="https://techcrunch.com/2026/07/14/openais-first-hardware-device-is-reportedly-a-screenless-speaker-that-can-move/">OpenAI&#8217;s first hardware is a screenless AI speaker that moves</a>. Bloomberg reports motorized parts, a rechargeable battery for room-to-room use, and a proactive &#8220;companion&#8221; persona built on ChatGPT; unveiling targeted for late 2026.</p></li><li><p><a href="https://techcrunch.com/2026/07/14/the-founder-of-hinge-raised-18m-to-build-a-new-ai-dating-service-overtone/">Hinge&#8217;s founder raises $18M for a voice-first AI dating app</a>. Justin McLeod&#8217;s Overtone replaces swiping with AI-curated introductions built around voice and audio; Match Group and FirstMark back it, with Esther Perel on the board.</p></li><li><p><a href="https://techcrunch.com/2026/07/14/spotify-expands-its-ai-push-with-a-chatgpt-like-music-assistant/">Spotify lets Premium users talk to the app to find music</a>. &#8220;Talk to Spotify&#8221; launches in beta for US, Ireland, and Sweden iOS/Android users 18+, picking songs and playlists conversationally from listening history.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/doordash-observeai-and-aws-partner-to-scale-customer-centric-ai-across-19-000-agents-302824595.html">DoorDash automates nearly 100% of support call reviews</a>. With Observe.AI and AWS, DoorDash now evaluates almost every customer interaction across 19,000 agents automatically, freeing QA teams for behavioral and safety analysis.</p></li><li><p><a href="https://www.nojitter.com/ai-automation/zoom-taps-into-big-market-for-ai-receptionists">Zoom moves into the AI receptionist market</a>. Zoom is building AI-driven virtual receptionist tools to automate front-desk call handling, competing for demand as businesses cut routine call-intake staffing.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/ai-gives-people-back-their-own-voice-chen-institute-and-science-prize-honors-neuroscientist-sergey-stavisky-302827712.html">A brain implant restores speech at 97.5% word accuracy</a>. Chen Institute and Science honor Sergey Stavisky&#8217;s neuroprosthesis, which now runs real-time voice synthesis with ~30ms delay for an ALS patient who has spoken 2.7M+ words with it.</p></li></ul><h2>&#129520; Platform News</h2><ul><li><p><a href="https://parallel.ai/blog/parallel-search-turbo">Cartesia and Parallel bring web search to voice agents at conversational speed</a>. Turbo mode extends Parallel&#8217;s low-latency search to Cartesia voice agents, hitting 220ms p50 search latency so agents can browse live results mid-conversation.</p></li><li><p><a href="https://developers.deepgram.com/changelog/voice-agent-changelog">Deepgram ships word-level timestamps in Flux</a>. Every word in the response now carries its own start and end time, tightening sync for captions, analytics, and downstream agent tooling.</p></li><li><p><a href="https://vapi.ai/blog/humanness-index-results">Vapi&#8217;s crowdsourced index puts Grok TTS closest to human speech</a>. After 11,000+ blind votes, Grok TTS scored 95 against a 100 human baseline, with MiniMax Speech 2.5 close behind at 93.</p></li><li><p><a href="https://www.hume.ai/blog/introducing-real-world-voiceeq-measuring-the-human-quality-of-voice-ai">Hume launches a benchmark for the &#8220;human quality&#8221; of voice AI</a>. Real World VoiceEQ scores 40+ models across 15+ dimensions using 1M+ human ratings, finding models have gotten better at speaking than at listening.</p></li><li><p><a href="https://gigazine.net/gsc_news/en/20260714-apple-speech-analyzer-benchmark/">Apple&#8217;s on-device SpeechAnalyzer beats Whisper Small on English</a>. Independent tests found a 2.12% word error rate on clear speech versus Whisper Small&#8217;s 3.74%, running fully on-device at roughly 3x the speed.</p></li></ul><h2>&#128214; Reading</h2><ul><li><p><a href="https://www.assemblyai.com/blog/how-to-catch-voice-agent-regressions-before-your-users-do">How to catch voice agent regressions before your users do</a>. AssemblyAI&#8217;s Griffin Sharp lays out proactive testing methods for catching production quality drops before they reach customers.</p></li><li><p><a href="https://livekit.com/blog/keeping-your-agent-conversation-on-track">Build a voice agent that won&#8217;t go off script</a>. LiveKit&#8217;s guide to keeping conversational agents on-topic during long sessions, with techniques for constraining scope and catching drift.</p></li><li><p><a href="https://www.resultsense.com/news/2026-07-17-clinical-ai-scribes-risks/">Five risks nobody&#8217;s tracking in clinical AI scribes</a>. A peer-reviewed study flags inconsistent consent, weak accented-speech performance, clinical background noise, missing human review, and unclear error accountability.</p></li><li><p><a href="/__u/sebastianbarros.substack.com/p/the-telco-ai-voice-is-bigger-than">Telcos are sitting on a voice AI opportunity bigger than cost savings</a>. Sebastian Barros argues operators should sell the primitives &#8212; phone numbers, fraud signals, deepfake detection, low-latency routing &#8212; every enterprise voice agent will need.</p></li><li><p><a href="https://www.coval.ai/blog/future-of-agentic-voice-webinar-recap">Coval recaps its &#8220;Future of Agentic Voice&#8221; webinar</a>. Replay and takeaways on where agentic voice AI is headed, covering trends and best practices for building and evaluating voice agents.</p></li></ul><h2><strong>&#128230; Releases</strong></h2><ul><li><p><span>LiveKit Agents: </span><a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.6">v1.6.6</a><span>.  Allow swapping models at runtime, more control for MCP tools and many improvements and fixes.</span></p></li><li><p><span>Pipecat: </span>No releases.</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.68">0.11.68</a>. Mistral Voxtral TTS, Gradium TTS and some other small improvements and fixes.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RealTime AI - Weekly Updates! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jul 13th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jul-13th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jul-13th-2026</guid><pubDate>Mon, 13 Jul 2026 19:12:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>&#128478;&#65039; Market and Product News</h2><ul><li><p><a href="https://techcrunch.com/2026/07/09/paris-based-ai-voice-startup-gradium-raises-100m-seed-backed-by-nvidia/">Gradium raises $100M seed backed by NVIDIA</a>. The Kyutai spinout expands its seed to $100M with NVIDIA, FirstMark, Eric Schmidt, and Xavier Niel; customers include Renault, and a Bay Area office is next.</p></li><li><p><a href="https://www.businesswire.com/news/home/20260707542838/en/Omilia-Powers-Taco-Bells-Expansion-of-Voice-AI-Across-890-U.S.-Drive-Thrus">Taco Bell scales Omilia voice AI to 890+ drive-thrus</a>. Live across 38 states, adapting to per-store menus and real-time stock; transaction times match or beat human order-taking, with higher staff retention.</p></li><li><p><a href="https://elevenlabs.io/blog/alpha-bank">Alpha Bank puts ElevenAgents in its call center</a>. One of Greece&#8217;s largest banks deploys a custom-voice agent handling first-contact service in Greek and English, with handoff to human advisors.</p></li><li><p><a href="https://pulse2.com/whispp-raises-e5-million-to-scale-real-time-on-device-voice-reconstruction-ai/">Whispp raises &#8364;5M for on-device voice reconstruction</a>. LUMO Labs leads; the app converts whispered or impaired speech into a clear natural voice in real time, fully on-device &#8212; no cloud, no added latency.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/study-finds-ai-voice-agents-increased-specialty-care-program-enrollment-rates-340-in-real-world-clinical-setting-302818900.html">Study Finds AI Voice Agents Increased Specialty Care Program Enrollment Rates 340% in Real-World Clinical Setting</a>. RadiantGraph/Oshi Health trial across 8,800 members: 9x outbound call volume, ~75 staff hours saved per 1,000 contacts, zero safety incidents.</p></li><li><p><a href="https://www.biometricupdate.com/202607/hydaway-introduces-real-time-enterprise-audio-deepfake-detection">Hydaway launches streaming audio deepfake detection</a>. RealityChek now flags synthetic speech in live calls by analyzing spectral patterns, prosody, and generative-model signatures &#8212; &#8220;a lie detector for live audio.&#8221;</p></li><li><p><a href="https://www.nojitter.com/contact-centers/5-numbers-showing-how-contact-centers-use-ai">5 numbers on how contact centers actually use AI</a>. 95% of companies run AI somewhere, but only 11% of CX leaders say it autonomously resolves significant interactions &#8212; and 46% are adding human oversight.</p></li></ul><h2>&#129520; Platform News</h2><ul><li><p><a href="https://openai.com/index/introducing-gpt-live/">OpenAI ships GPT-Live, full-duplex voice for ChatGPT</a>. The model listens while it speaks &#8212; backchannels, interruptions, knowing when to stay quiet &#8212; and delegates hard queries to a frontier model. No API yet.</p></li><li><p><a href="https://x.com/OpenAIDevs/status/2074255420831735824">OpenAI updates Realtime: GPT-Realtime-2.1 and a 25% p95 latency cut</a>. New Realtime models improve alphanumeric recognition and noise handling; improved caching cuts p95 latency at least 25% across Realtime voice models.</p></li><li><p><a href="https://www.cartesia.ai/blog/ink-2">Cartesia releases Ink-2 streaming STT</a>. Ranked #1 lowest WER on Artificial Analysis streaming; 8% WER on 14-accent call-center audio, semantic endpointing, 0.1s time-to-final. On LiveKit, Vapi, Pipecat.</p></li><li><p><a href="https://www.assemblyai.com/blog/universal-3-5-pro-async">AssemblyAI launches Universal-3.5 Pro for async at $0.21/hr</a>. Native code-switching across 18 languages, its best speaker diarization yet, and contextual prompting; beats Nova-3 Multilingual and Scribe v2 on code-switching WER.</p></li><li><p><a href="https://x.ai/news/new-flagship-voices">xAI adds 21 multilingual flagship voices to Grok</a>. Every voice is natively multilingual across 25+ languages, available in the Voice Agent and TTS APIs, with voice cloning from a 120-second clip (US-only).</p></li><li><p><a href="https://www.businesswire.com/news/home/20260708009575/en/Omilia-Launches-the-Only-Native-Voice-in-Enterprise-CX">Omilia launches Lexis, TTS built into its CX platform</a>. Sub-45ms first-audio latency with no third-party API round-trips; sentence-level tone and pacing analysis plus branded voice cloning from a short recording.</p></li><li><p><a href="https://livekit.com/blog/async-tools-voice-agents">LiveKit ships async tools for voice agents</a>. Agents keep talking while slow tools run: streamed progress, filler audio, mid-task cancellation, and duplicate-call protection. A 15s backend op gets acknowledged in ~1s.</p></li><li><p><a href="https://www.retellai.com/blog/ios-call-screening">Retell handles iOS 26 call screening automatically</a>. Voice agents detect Apple&#8217;s screening prompt and deliver talk tracks tuned to the ~30-second window before the human picks up &#8212; relevant to ~150M US iPhones.</p></li></ul><h2>&#128214; Reading</h2><ul><li><p><a href="https://hackernoon.com/we-used-benchmaxxers-favourite-trick-to-climb-10-places-on-the-hugging-face-open-asr-leaderboard">Speechmatics: how we benchmaxxed the Open ASR Leaderboard</a>. Fine-tuning on AMI-like data jumped them 10 places (6.90 &#8594; 6.38 WER) &#8212; and a candid case for how gameable public ASR leaderboards really are.</p></li><li><p><a href="https://livekit.com/blog/your-model-isnt-bad-at-tool-calling">Your model isn&#8217;t bad at tool calling &#8212; your serving stack is</a>. When the provider lacks a parser for the model&#8217;s tool-call format, calls leak into text and get read aloud by TTS. Includes a one-line curl diagnostic.</p></li><li><p><a href="https://www.assemblyai.com/blog/conversation-context-voice-agents">How conversation context fixes STT&#8217;s worst failure modes</a>. Feeding both sides of the dialog to the model fixes emails, names, and one-word replies; up to 100 turns of carryover plus a ~1,500-char agent context.</p></li></ul><h2><strong>&#128230; Releases</strong></h2><ul><li><p><span>LiveKit Agents: </span><a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.5">v1.6.5</a><span>. Many improvements and fixes, including conversation aware STT for better accuracy and expressive mode for TTS providers.</span></p></li><li><p><span>Pipecat: </span>v1.5.0. Pipecat Flows is now officially part of Pipecat, Together AI STT + TTS, text normalization before synthesis and tons of stability fixes across turn management, TTS/STT, and transports.</p></li><li><p>TEN Framework: No releases.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RealTime AI - Weekly Updates! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jun 29nd 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jun-29nd-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jun-29nd-2026</guid><pubDate>Tue, 30 Jun 2026 15:07:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>&#128478;&#65039; Market and Product News</h2><ul><li><p><a href="https://www.fiercehealthcare.com/ai-and-machine-learning/assort-health-scores-120m-series-c-scale-voice-ai-agent-platform-healthcare">Assort Health raises $120M Series C at $1.2B valuation</a>.The voice AI agent platform for healthcare reports 20x revenue growth in 15 months and 115% increase in labor capacity across 190M specialty patient interactions.</p></li><li><p><a href="https://www.globenewswire.com/news-release/2026/06/22/3315343/0/en/prosper-ai-raises-30m-from-andreessen-horowitz-to-scale-the-first-ai-platform-to-run-the-entire-patient-journey.html">Prosper AI raises $30M Series A from a16z</a>. Platform automates the full patient journey &#8212; scheduling through billing &#8212; and reports 5x revenue growth in six months while winning 80% of competitive RFPs.</p></li><li><p><a href="https://www.coval.ai/blog/coval-series-a">Coval raises $28M Series A to fix voice agent reliability gap</a>. The voice AI testing platform addresses the fact that ~95% of agents work in demos but only ~62% survive their first week live.</p></li><li><p><a href="https://www.businesswire.com/news/home/20260624276012/en/Kotoba-Technologies-Raises-$10-Million-in-Seed-Funding-to-Expand-Real-Time-Voice-AI-Platform-Across-East-Asia">Kotoba Technologies raises $10M seed for sub-50ms East Asian voice translation</a>. Kindred Ventures leads, with Salesforce Ventures and Sony Innovation Fund; the &#8220;Koto&#8221; model handles Japanese, Korean, and Chinese speech-to-speech translation on a single smartphone.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/valence-ai-raises-5-million-secures-us-patents-on-real-time-emotional-detection-from-live-speech-302808293.html">Valence AI raises $5M seed, secures two patents on real-time emotional detection</a>. Differential Ventures leads; Pulse Emotion model reports 92% accuracy and 30% handle-time reduction in production contact center deployments.</p></li><li><p><a href="https://techcrunch.com/2026/06/15/salesforce-acquires-ai-customer-service-platform-fin-for-3-6b/">Salesforce acquires Fin for $3.6B, folds into Agentforce</a>. Fin&#8217;s Apex AI resolves 76% of queries autonomously across voice, chat, and messaging for 30,000+ enterprise customers; deal closes Q4 FY2027.</p></li><li><p><a href="https://arrowhead.ai/case-studies/tata-1mg">Arrowhead.ai beats human conversion rates for Tata 1MG abandoned cart calls</a><br>Voice AI outperformed human agents by 15% on abandoned cart recovery for Tata 1MG, with up to 45% higher conversions across BFSI deployments.</p></li></ul><h2>&#129520; Platform News</h2><ul><li><p><a href="https://krisp.ai/blog/krisp-anounces-voice-security-speech-analytics/">Krisp launches Voice Security for real-time deepfake detection in contact centers</a>. Detects cloned voices in real time before account access; deepfakes have grown 22x in three years and human agents catch them only ~60% of the time.</p></li><li><p><a href="https://cryptobriefing.com/openai-chatgpt-bidi-1-voice-model/">OpenAI tests GPT-Bidi-1 full-duplex voice model internally</a>. Spotted in ChatGPT code in mid-June; the model listens and speaks simultaneously, supporting mid-sentence adjustments and interruptions without freezing context.</p></li><li><p><a href="https://www.marktechpost.com/2026/06/24/gradium-launches-stt-translate-and-s2s-translate-real-time-speech-translation-models-beating-gpt-realtime-translate-on-accuracy-and-latency/">Gradium launches STT-Translate and S2S-Translate, beating GPT Realtime on BLEU and latency</a>. Two-model pipeline covers 5 languages and 20 pairs via a single WebSocket; s2s-translate averages 3.0s vs. GPT Realtime&#8217;s 3.6s, with voice cloning support.</p></li><li><p><a href="https://www.speechmatics.com/company/articles-and-news/introducing-melia-multilingual-speech-to-text-model">Speechmatics releases Melia: 55+ language multilingual STT with auto language tagging</a>. Single model handles all supported languages with automatic code-switching detection; beats Deepgram and Microsoft on FLEURS; priced from $0.129/hr.</p></li><li><p><a href="https://www.retellai.com/blog/introducing-conductor">Retell AI launches Conductor: AI builder for voice agent contact centers</a><br>Describes agents in plain English, catches problems before production, and runs simulations; internal data shows 2.5x faster to production, 60% less build time.</p></li></ul><h2>&#128214; Reading</h2><ul><li><p><a href="https://lemonslice.com/blog/connection-time">LemonSlice cuts avatar connection time in half: median drops from 5.4s to 2.9s</a>. P99 falls from 37.8s to 9.0s (76% improvement) by proactively warming the VAE, redefining GPU-ready criteria, and parallelizing WebRTC setup.</p></li><li><p><a href="https://vllm.ai/blog/2026-06-23-vllm-omni-tts">vLLM-Omni publishes TTS engineering guide: Qwen3-TTS, Fish Speech S2 Pro, Higgs Audio V3</a>. Covers serving bottlenecks unique to TTS autoregression; Qwen3-TTS gets 61.5% audio throughput gains, VoxCPM2 gets 172% via specialized kernels.</p></li><li><p><a href="https://voiceaiandvoiceagents.com/">Voice AI Guide 2026 + Smart Turn v3: open-source turn detection for 23 languages</a>. Pipecat&#8217;s updated primer declares turn detection &#8220;basically solved&#8221; &#8212; hybrid VAD + Smart Turn + LLM tagging; Smart Turn v3 runs in 12ms CPU, covers 23 languages.</p></li><li><p><a href="https://x.com/MiravoiceAI/status/2069847391285486081">Miravoice&#8217;s 100k-call voice comparison: Rime vs. ElevenLabs vs. Google</a>. Real outbound calls to US households; among callers who stayed past the intro, Rime had the highest survey completion rate &#8212; voice was the variable that moved the needle.</p></li><li><p><a href="https://aws.amazon.com/blogs/machine-learning/build-a-healthcare-appointment-agent-with-amazon-nova-2-sonic/">AWS tutorial: healthcare appointment agent with Amazon Nova 2 Sonic + AgentCore</a>. End-to-end reference build for outbound appointment reminders with voice auth, schedule management, and mid-call language switching; no-show rates average 5&#8211;30% in US healthcare.</p></li></ul><h2><strong>&#128230; Releases</strong></h2><ul><li><p><span>LiveKit Agents: </span><a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.3">v1.6.3</a><span> - </span><a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.4">v1.6.4</a><span>. Small releases adding Protoface avatar plugin and some fixes.</span></p></li><li><p><span>Pipecat: </span>No releases.</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.67">v0.11.67</a>. Small release adding Spatius avatar plugin and some fixes.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RealTime AI - Weekly Updates! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jun 22nd 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jun-22nd-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jun-22nd-2026</guid><pubDate>Mon, 22 Jun 2026 17:45:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://www.prnewswire.com/news-releases/bland-surpasses-100m-funding-with-new-series-c-to-advance-voice-ai-for-complex-high-stakes-conversations-302801583.html">Bland raises $100M Series C led by Dell Technologies Capital</a>. Processes 3.5M calls/week and 175M calls last year; $100M+ raised total. Bets on proprietary voice models &#8212; no OpenAI or Anthropic underneath.</p></li><li><p><a href="https://techcrunch.com/2026/06/17/deepl-acquires-mixhalo-for-live-event-audio-streaming-and-translation/">DeepL acquires Mixhalo for real-time live audio</a>. Ultra-low-latency stadium audio infrastructure folds into DeepL Voice, targeting live events, conferences, and Amazon Connect. Opens DeepL&#8217;s first SF office.</p></li><li><p><a href="https://elevenlabs.io/blog/poland-invests-in-elevenlabs">Poland takes a government stake in ElevenLabs</a>. Vinci (part of BGK Group) joins a16z, Sequoia, and ICONIQ as shareholders. Launches AI Lab Poland with LOT Polish Airlines, InPost, and healthcare partners.</p></li><li><p><a href="https://simplertc.com/">Sesame acquires SMPL</a>. SMPL brings codec, DSP, and echo cancellation expertise to Sesame&#8217;s conversational AI agents. Brendan Iribe: &#8220;some of the best in audio.&#8221;</p></li><li><p><a href="https://blog.google/products-and-platforms/devices/google-nest/google-home-speaker-gemini-features/">Google Home Speaker opens preorders at $99.99, ships June 25</a>. First device built for Gemini for Home. Runs local models for noise cancellation and sound separation. Gemini Live handles open-ended conversation without a wakeword.</p></li><li><p><a href="https://humannessindex.vapi.ai/">Vapi launches Humanness Index: xAI Grok TTS leads at 96/100</a>. Crowdsourced blind leaderboard ranking 21 TTS models against a human baseline (100). 9,350+ votes; ElevenLabs v3 scores 93, MiniMax Speech 2.5 scores 92.</p></li><li><p><a href="https://www.prnewswire.com/apac/news-releases/tencent-cloud-and-inworld-ai-announce-strategic-partnership-to-deliver-a-one-stop-lifelike-realtime-voice-ai-solution-302799015.html">Tencent Cloud and Inworld AI announce real-time voice partnership</a> &#183; Jun 16, 2026<br>Inworld TTS (sub-130ms first-chunk latency, 100+ languages) integrates into Tencent RTC&#8217;s 3,200-node global infrastructure. Developers pick Inworld TTS directly from the Tencent RTC console.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://www.cartesia.ai/launch/">Cartesia ships Sonic-3.5 (TTS) and Ink-2 (STT)</a>. Both rank #1 on Artificial Analysis streaming leaderboards &#8212; sub-90ms TTS latency, 100ms STT. Now the only provider simultaneously leading both sides of the voice stack.</p></li><li><p><a href="https://soniox.com/blog/soniox-v5-real-time">Soniox v5 Real-Time STT model released</a>. Adds real-time translation, code-switching across 60+ languages, and faster semantic endpointing. v4 retires June 30 with automatic routing to v5 after that.</p></li><li><p><a href="https://deepgram.com/learn/deepgram-australia-endpoint-now-generally-available">Deepgram Australia endpoint is generally available</a>.</p></li><li><p><a href="https://www.medianama.com/2026/06/223-gnani-ai-prisma-v2-5-speech-recognition-model-better-accuracy-sarvam/">Gnani AI launches Prisma v2.5 STT: #1 in 8 of 9 Indian language benchmarks</a>. </p></li><li><p><a href="https://mastra.ai/blog/mastra-inworld-realtime-voice">Inworld Realtime API now integrates with Mastra and Voximplant</a>. One WebSocket session gives Mastra agents speech I/O, semantic VAD, barge-in, and tool calling via @mastra/voice-inworld-realtime. Voximplant adds the same for phone calls, SIP, and WhatsApp.</p></li><li><p><a href="https://livekit.com/blog/solving-end-of-turn-detection">LiveKit Turn Detector v1.0: 9.9% false-cutoff rate at 300ms</a>. Fuses semantic (audio-to-LLM embeddings) and acoustic (intonation, pitch) branches to detect turn end without reading transcripts. Beats Deepgram Flux (12.9%) and ultraVAD (27.7%). Supports 14 languages. Free on LiveKit Cloud; v1-mini ships in Agents SDKs (Python 1.6.1, TypeScript 1.4.7) under Apache-2.0.</p></li><li><p><a href="https://livekit.com/benchmarks/eot-bench">LiveKit releases eot-bench: open end-of-turn benchmark</a>. Public leaderboard and open dataset for comparing end-of-turn models on shared ground truth. GitHub harness at livekit/eot-bench; anyone can submit results.</p></li><li><p><a href="https://blogs.nvidia.com/blog/nvidia-xr-ai/">NVIDIA XR AI enters public beta for AR glasses</a>. Developer library for real-time agentic apps on XR devices. Agents perceive via video, audio, and depth sensors; connect to enterprise knowledge; and deliver hands-free guidance for manufacturing and surgery.</p></li></ul><h2>&#128214; Reading</h2><ul><li><p><a href="https://www.nojitter.com/contact-centers/what-the-landmark-google-ruling-means-for-contact-center-ai">What the Google liability ruling means for contact center AI.</a> A German court treated AI Overviews output as &#8220;direct corporate speech,&#8221; making Google liable for hallucinations. Applies directly to enterprise voice agents making policy promises.</p></li><li><p><a href="https://webrtc.ventures/2026/06/slug-voice-ai-security-webrtc-livekit-guardrails/">Voice AI Security: Building Realtime Voice Agents with WebRTC, LiveKit, and Sensitive Data Guardrails</a> (webrtc.ventures)<span>. A reference architecture for secure realtime voice AI, show how guardrails are enforced inside a live LiveKit pipeline, and cover the layered security model that production systems need.</span></p></li><li><p><a href="https://www.assemblyai.com/blog/speech-to-speech-for-voice-agents">Speech-to-speech for voice agents: cascaded vs. end-to-end</a> (AssemblyAI). Cascaded (STT+LLM+TTS) still dominates production because of observability and flexibility; end-to-end models lack production-grade accuracy yet. Covers what sub-1s latency requires at each stage.</p></li><li><p><a href="https://vapi.ai/blog/designing-conversations">Vapi: designing conversations for the ear</a> (VAPI). Six principles for converting chat scripts to voice flows: single-idea turns, binary choices over open questions, immediate confirmation. One staffing company cut call duration 40s by changing a single question type.</p></li><li><p><a href="https://www.coval.ai/blog/voice-agent-vendor-testing">Voice Agent Vendor Testing: How to Run a Bake-Off</a> (Coval). Guide to rigorous platform evaluations: scenario design, scoring rubrics, and why comparing vendors on different call sets or metrics produces useless results.</p></li><li><p>[Youtube] <a href="https://www.youtube.com/watch?v=tlqcNWI7xy8">Kwindla Kramer on building Pipecat and the bet nobody else wanted to take</a> (Bluejay)  Ten of twelve 2016 investors passed; NVIDIA and AWS standardize on Pipecat today. His call: open, low-latency specialized models replace big-lab APIs in production within 2&#8211;3 years.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.1">v1.6.1</a> - <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.2">v1.6.2</a>. Turn Detector v1.0, Assembly AI Universal 3.5, Gemini 3.1. Flash TTS, Soniox STT v5 and many bufixes and improvements.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.4.0">v1.4.0</a>. New framework Pipecat Evals included!, more advanced support for real-time models, Added local Moonshine STT and the new AIC Quail VAD analyzer, Simplified Tools registration and many other small improvements. </p></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jun 15th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jun-15th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jun-15th-2026</guid><pubDate>Mon, 15 Jun 2026 20:24:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://elevenlabs.io/blog/uk-mou-and-expansion">ElevenLabs signs MOU with UK government, triples London HQ</a>. Partnership with the UK&#8217;s Department for Science, Innovation and Technology covers accessibility, Welsh-language AI, and security research with the AI Security Institute; London headcount doubles to 200 this year.</p></li><li><p><a href="https://www.cxtoday.com/contact-center/why-voice-ai-adoption-is-accelerating-in-2026/">Voice AI crosses the enterprise threshold in contact centers</a>. Gartner projects 70% of support interactions automated by end-2027; AudioCodes reported 50%+ YoY growth in Q1 2026; post-call summaries and agent assist are where ROI is materializing today.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/">Google Gemini 3.5 Live Translate lands in 70+ languages</a>. Near real-time speech-to-speech translation that preserves speaker intonation, pacing, and pitch; available via Gemini Live API for developers, private preview in Google Meet, and global rollout on Google Translate Android/iOS.</p></li><li><p><a href="https://krisp.ai/blog/krisp-launches-v3-real-time-voice-translation/">Krisp launches Voice Translation v3 with 96% accuracy across 61 languages</a>. 93&#8211;97% accuracy range across 30 benchmarked domains; 90% of multilingual healthcare calls completed without a human interpreter; self-serve developer API now live with JS and Python SDKs.</p></li><li><p><a href="https://www.gladia.io/blog/solaria-3-speech-to-text-model-for-european-languages">Gladia launches Solaria-3, top ASR on European production audio</a>. Solaria-3 ranks #1 at 6.4% WER on financial and business speech; optimized for English, French, German, Spanish, and Italian.</p></li><li><p><a href="https://www.misolabs.ai/blog/miso-tts-8b">Miso Labs open-sources Miso One TTS model</a>. 8B-parameter emotive TTS built on the Sesame CSM architecture; one-shot voice cloning from an audio prompt; 110ms latency; weights on Hugging Face under a modified MIT license.</p></li><li><p><a href="https://gradium.ai/blog/gradium-tts-upgrade">Gradium TTS now reads email addresses correctly 97% of the time</a>. Updated default model tops the field on email addresses (97%), phone numbers, and time expressions.</p></li><li><p><a href="https://www.resemble.ai/resources/chatterbox-multilingual-v3-tts-with-embedded-watermarking-for-25-languages">Resemble AI ships Chatterbox TTS Multilingual v3 in 25 languages</a>. Next general-purpose TTS model in the Chatterbox family. It supports 25 total languages, and ships meaningful improvements in speaker similarity, hallucination rate, and conversational naturalness.</p></li><li><p><a href="https://kyutai.org/blog/2026-06-10-interactivity">Kyutai applies RL post-training to sharpen full-duplex speech interactivity</a>. Post-training alignment for full-duplex models like Moshi using RL reward functions targeting turn-taking, barge-in, and conversational responsiveness.</p></li><li><p><a href="https://inworld.ai/blog/consumer-ai-cost-pricing">Inworld cuts voice AI API prices by ~50%</a>.  TTS-2 drops to ~$10/1M chars, STT to $0.10/hr on-demand; framed as fixing the math for consumer AI apps where 97% of users never pay.</p></li><li><p><a href="https://community.livekit.io/t/introducing-livekit-portal-production-grade-stack-for-teleoperation-and-remote-inference-on-robots/1412">LiveKit Portal ships production-grade teleoperation stack for robots</a>. Time-aligned observations, multi-operator support, live pipeline latency metrics, and network-agnostic deployment &#8212; built on LiveKit&#8217;s realtime infrastructure for remote inference on robots.</p></li><li><p><a href="https://ai-coustics.com/tyto">ai-coustics launches Tyto to predict voice agent audio failures in real time</a>. Scores incoming call audio across 7 acoustic dimensions to predict VAD, ASR, and speech-to-speech failures before they happen; deployed by PolyAI, telli, and LiveKit.</p></li><li><p><a href="https://www.linkedin.com/posts/retellai_day-1-of-5-retell-launch-week-2026-activity-7469792923518160896-LGXt">Retell Launch Week 2026</a>.  Live Monitoring, Built-in CRM, the context layer for voice agents, with two-way real-time syncing to Salesforce and HubSpot, Colloquial Model and Custom Dashboards.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://www.linkedin.com/posts/imbert-hugo_you-cant-compute-wer-on-production-audio-ugcPost-7465236626529538048-gAyq">Why you can&#8217;t compute WER on production audio &#8212; and what to do instead</a>. Explains why voice teams can&#8217;t label production data for WER, then introduces Noisekit &#8212; an open-source tool that degrades clean datasets to simulate real-world STT conditions.</p></li><li><p><a href="https://arxiv.org/abs/2606.10231">Mel-LLM: encoder-free speech LLM that processes Mel spectrograms directly</a>. Microsoft team feeds Mel spectrogram patches via linear projection into an LLM &#8212; no speech encoder; competitive ASR results, especially when initialized from multimodal checkpoints like Phi-4-MM.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.6.0">v1.6.0</a>. Introducing asynchronous tools, new filler phrases support, provider updates for Deepgram, ElevenLabs, Google, AWS, Sarvam and others and many bufixes and improvements.</p></li><li><p>Pipecat: No releases.</p></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jun 8th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jun-8th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jun-8th-2026</guid><pubDate>Mon, 08 Jun 2026 20:59:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://techcrunch.com/2026/06/03/these-two-founders-left-goldman-and-meta-to-build-voice-ai-for-markets-everyone-else-overlooked/">AethexAI raises $3M to build voice AI for Africa and the Middle East</a>. Ex-Goldman and Meta founders built Kora models (300M&#8211;1.7B params) for Arabic, French, and English dialects; already handling 17K calls/day for debt collection, KYC, and telecoms.</p></li><li><p><a href="https://elevenlabs.io/blog/lot-polish-airlines-announcement">LOT Polish Airlines deploys ElevenLabs voice agents for customer service</a>. Poland&#8217;s flag carrier becomes one of the first major European airlines to run ElevenLabs-powered voice agents across customer interactions.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://www.marktechpost.com/2026/06/06/nvidia-releases-nemotron-3-5-asr-a-600m-parameter-cache-aware-streaming-model-transcribing-40-language-locales-in-real-time/">NVIDIA Nemotron 3.5 ASR transcribes 40 languages locally under 100ms</a>. NVIDIA released a new multilingual speech-to-text model today: Nemotron 3.5 ASR, ideal for voice agents. Available in both multilingual (Nemotron 3.5 ASR) and English-only (Nemotron 3 ASR) checkpoints. It&#8217;s the lowest-latency STT model we&#8217;ve tested. It&#8217;s also completely open source, fine-tunable, and you can host it on your own infrastructure.</p></li><li><p><a href="https://github.com/FluidInference/FluidAudio">NVIDIA's Nemotron 3.5 ASR Streaming Multilingual available for Apple Silicon</a>. Model is now available through FluidAudio CoreML optimized for Apple Silicon so apps can run ~40-language real-time ASR entirely on device, no cloud required.</p></li><li><p><a href="https://livekit.com/blog/livekit-cpp-sdk-official-release">LiveKit ships C++ SDK 1.0.0 for robotics and embedded systems</a>. Native C++ client for realtime audio, video, and data tracks. Runs on Linux, macOS, Windows; ARM targets include NVIDIA Jetson, Raspberry Pi, and Rockchip. Hardware encoder acceleration included; ROS2 bridge on the roadmap.</p></li><li><p><a href="https://x.com/HappyRobot/status/2062573186714333469">HappyRobot launches its own TTS model</a>. Built for low-latency deployment with accurate pronunciation of numbers, codes, and alphanumerics &#8212; the edge cases that break generic TTS in freight and logistics voice agents.</p></li><li><p><a href="https://microsoft.ai/news/mai-voice-2/">Microsoft launches MAI-Voice-2: zero-shot voice cloning across 17 languages</a>. Clones a voice from 5&#8211;60s of audio with no retraining; wins 72% in head-to-head preference tests vs MAI-Voice-1; code-switches in Hindi-English and Spanish-English. Available in Foundry, VSCode, and Dynamics 365 Contact Center.</p></li><li><p><a href="https://x.com/AodenTeoMT/status/2062204362102100295">Miso Labs open-sources Miso One: 8B TTS at 110ms with human-level emotional prosody</a>. 8-billion-parameter open-source TTS model for highly expressive speech; 110ms latency; According to them &#8220;the most emotive voice model in the world&#8221;.</p></li><li><p><a href="https://github.com/rednote-hilab/dots.tts">Rednote open-sources dots.tts: first fully continuous TTS pipeline (no codec)</a>. 2B-parameter end-to-end autoregressive TTS with no discrete tokens anywhere &#8212; continuous AudioVAE at 48kHz feeding a flow-matching acoustic head. 24 languages, zero-shot voice cloning, Apache 2.0.</p></li><li><p><a href="https://magenta.withgoogle.com/magenta-realtime-2">Google Magenta RealTime 2</a>: open-weights real-time music generation with text, audio, and MIDI. 230M and 2.4B model sizes; streams audio from text prompts or note input under 200ms. Apache 2.0; community PyTorch port with ZeroGPU demos appeared within hours of release.</p></li><li><p><a href="https://x.com/TencentRTC/status/2061703470802190715">Tencent RTC and Soniox partner for enterprise voice AI</a>. Combines Soniox&#8217;s STT 60+ language accuracy with Tencent&#8217;s 3,200-node global network to deliver under 300ms voice AI latency in 200+ countries.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://www.coval.ai/blog/best-text-to-speech-providers-in-2026-how-to-choose-(and-why-vendor-benchmarks-lie)/">Best TTS Providers 2026: Why Vendor Benchmarks Lie</a> (Coval). Coval benchmarks ElevenLabs, Cartesia, OpenAI, Deepgram, and 10 others; finds latency is no longer the top differentiator &#8212; emotional control, multilingual depth, and cost now separate the leaders.</p></li><li><p><a href="https://www.coval.ai/blog/best-speech-to-text-providers-in-2026-independent-benchmarks-and-how-to-choose">Best STT Providers 2026: Independent Benchmarks &amp; How to Choose</a> (Coval). Accuracy on clean English has plateaued; the 14-provider comparison finds 30&#215; price spread across the market and end-of-turn detection speed as the new differentiator in voice agent workloads.</p></li><li><p><a href="https://medium.com/p/barge-in-and-full-duplex-the-architecture-that-makes-voice-agents-feel-human-48ef1b7b7fa6">Barge-in and full-duplex: the architecture that makes voice agents feel human</a>. Walks through the five-step interruption pipeline (VAD &#8594; intent classification &#8594; TTS cancel &#8594; LLM abort &#8594; re-listen) and explains why event-driven decoupled design is required to stay under 300ms.</p></li><li><p><a href="https://www.assemblyai.com/blog/streaming-speaker-diarization">Streaming speaker diarization: How to identify who&#8217;s speaking in real time</a> (AssemblyAI). Streaming speaker diarization identifies who is speaking in real time with low-latency labels. Learn how it works and when to use it for live apps.</p></li><li><p><a href="https://livekit.com/blog/mongodb-voice-agent-memory">Building voice agent persistent memory with MongoDB Atlas Vector Search</a> (LiveKit) LiveKit tutorial showing RAG + hybrid rankFusion recall for cross-session voice agent memory; user profile loads via vector search before the first word of the conversation.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.17">v1.5.17</a>. Adding more model options for LiveKit inference and many fixes and improvements.</p></li><li><p>Pipecat:No releases.</p></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Jun 1st 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-jun-1st-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-jun-1st-2026</guid><pubDate>Tue, 02 Jun 2026 20:48:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://thenextweb.com/news/parloa-turns-its-350-million-war-chest-into-a-partnership-web-spanning-sap-microsoft-and-openai">Parloa announced strategic partnerships with SAP, Microsoft, OpenAI, Five9, and Epic</a>. The Berlin-founded AI agent management platform valued at $3 billion, has announced strategic partnerships five months after raising $350 million in its Series D.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/greenhouse-completes-acquisition-of-ezra-ai-labs-bringing-conversational-ai-to-the-hiring-process-302782372.html">Greenhouse folds voice AI interviewing into its hiring platform with Ezra AI Labs acquisition</a>. Acquires Ezra AI Labs to automate initial candidate screening via structured voice conversations; applications per recruiter on Greenhouse are up 412% since 2023.</p></li><li><p><a href="https://news.gtp.gr/2026/05/28/greece-explores-ai-voice-technology-for-tourism-platforms-and-digital-tours/">Greece Explores AI Voice Technology with ElevenLabs for Tourism Platforms</a>. Will add voice to gov.gr and Visit Greece; accessibility focus targets elderly and disabled citizens interacting with state services by voice, plus 122-language museum tours.</p></li><li><p><a href="https://techcrunch.com/2026/05/28/sesame-the-conversational-ai-startup-from-oculus-founders-launches-its-ios-app/">Sesame launches iOS app with four persistent personal voice agents</a>. Maya, Miles, Simone, and Charlie carry persistent memory and run parallel web searches mid-conversation. From Oculus founders; available in 39 countries, free.</p></li><li><p><a href="https://technode.com/2026/05/29/iflytek-launches-40g-ai-glasses-with-glassclaw-ai-agent-and-advanced-noise-recognition/">iFlytek ships 40g AI glasses with GlassClaw agent and 122-language translation</a>. GlassClaw handles meeting transcription, email, and workflow execution; world-first lip movement recognition noise reduction</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai">NVIDIA launches Cosmos 3, the first open omnimodel generating video, audio, and robot action trajectories</a>. Trained on 20T tokens including 400M videos and ambient sound; outputs text, images, video, ambient audio, and action trajectories. Two variants: Super and Nano, with Edge coming.</p></li><li><p><a href="https://avtr-1.avaturn.live/">Avaturn releases AVTR-1: open-weights real-time conversational avatar model from a single photo</a>. Generates full-face 512&#215;512 video at 25fps; jointly models speaking and listening behavior, unlike prior portrait-animation systems. Free for commercial use under $10M ARR.</p></li><li><p><a href="https://slator.com/alibaba-speech-translation-model-triples-language-coverage/">Alibaba Updates Speech Translation Model</a><strong>. </strong>Alibaba&#8217;s Qwen team launched Qwen3.5-LiveTranslate-Flash, expanding its AI live speech translation model from 18 to 60 languages and adding real-time voice cloning.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p>[Youtube] <a href="https://elevenlabs.io/events/elevenlabs-summit/waw-26/resource/keynote-war-2026">ElevenLabs Warsaw Summit keynote: on-device voice AI and LOT Airlines case study</a> (ElevenLabs). Mati Staniszewski demos a new on-device architecture delivering cloud-quality voice AI fully offline on consumer hardware; LOT Polish Airlines shares production deployment results.</p></li><li><p><a href="https://decrypt.co/369042/inaudible-audio-attacks-hijack-ai-voice-models">Inaudible audio commands can hijack AI voice models with up to 96% success rate</a>. Zhejiang University embeds imperceptible commands in audio; attacks transferred from 13 open models to commercial systems from Microsoft Azure and Mistral with high reliability.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.13">v1.5.13</a>-<a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.15">v1.5.15</a>. GPT RealTime 2 model, GnaniAI STT, makes Google Model Garden LLMs available to Agents, Cartesia Ink-2 STT, Respecheer TTS, Fading support for background audio, and tons of fixes and improvements.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.3.0">v1.3.0</a>. This release brings the power of Pipecat Subagents directly into Pipecat itself, making multi-agent systems a first-class part of the framework.  New reasoning model support, Massive Smart Turn startup time and memory improvements, New turn-based STT support, New Vonage WebRTC transport.  Also Whisker 2.0.0 is now available!</p></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - May 25th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-may-25th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-may-25th-2026</guid><pubDate>Tue, 26 May 2026 17:36:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://elevenlabs.io/blog/bringing-voice-ai-into-the-classroom-with-elevenlabs">ElevenLabs releases Albert Einstein voice agent built from his written archives</a>. The agent recreates Einstein&#8217;s voice from historical recordings. Users can converse in real time with a model grounded in his actual written work.</p></li><li><p><a href="https://www.nojitter.com/ai-voice/quiq-now-supports-voice-ai">Quiq adds Voice AI to its agentic customer service platform</a>. Quiq&#8217;s platform now handles voice alongside digital channels, with AI agents that match tier-one human agent reasoning and pass full conversation context on human escalation.</p></li><li><p><a href="https://www.prweb.com/releases/palabraai-real-time-ai-voice-translator-hits-1m-arr--grows-17x-in-six-months-302779808.html">Palabra.ai hits $1M ARR -- Grows 17x in Six Months</a>. Palabra.ai (Real-Time AI Voice Translator) now translates thousands of meetings, webinars, and live broadcasts every month, in 60+ languages and the original speaker's voice.</p></li><li><p><a href="https://investors.soundhound.com/news-releases/news-release-details/soundhound-ai-acquire-liveperson-combining-proprietary-voice">SoundHound AI To Acquire LivePerson</a>. The combination unifies SoundHound&#8217;s industry-leading voice and agentic AI platform with LivePerson&#8217;s digital engagement capabilities.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://x.com/ArtificialAnlys/status/2057878247782908109">Cartesia Sonic 3.5 hits #1 on Speech Arena</a>. Cartesia&#8217;s Sonic 3.5 took the top spot on the Artificial Analysis Speech Arena, beating Inworld and Gemini. It hits 82ms end-to-end time-to-first-audio with support for 42 languages.</p></li><li><p><a href="https://www.prnewswire.com/apac/news-releases/tencent-cloud-and-stream-partner-to-accelerate-the-development-of-real-time-multimodal-ai-agents-302774565.html">Tencent RTC and Stream partner to bring Tencent network to Stream framework</a>. Tencent Real-Time Communication is now an official transport plugin for Stream&#8217;s open-source Vision Agents framework, giving developers access to Tencent&#8217;s 3,200-node backbone with sub-300ms global latency.</p></li><li><p><a href="https://rime.ai/coda">Rime launches Coda, a new flagship TTS model with 180+ voices for enterprise</a>. Coda is Rime&#8217;s new flagship model built specifically for enterprise customer interactions &#8212; 180+ voices, brand voice cloning, SOC 2 Type II and HIPAA certified, available on-premises or in the cloud.</p></li><li><p><a href="https://elevenlabs.io/speech-engine">ElevenLabs launches Speech Engine to turn any chat agent into a voice agent</a>. Speech Engine combines ElevenLabs&#8217; STT, TTS, and voice orchestration into a single pipeline to add human-like voice to your existing chat agent with a single prompt. </p></li><li><p><a href="https://livekit.com/blog/ship-voice-agent-on-any-website-script-tag">LiveKit launches voice agent widget deployable on any website with a single script tag</a>. The widget supports voice, video, screen share, and text chat, with per-user context passing and domain locking.</p></li><li><p><a href="https://docs.livekit.io/agents/models/avatar/plugins/runway/">LiveKit adds RunwayML Characters support for real-time AI avatars from a single image</a>. Runway provides the animated avatar video layer with eye contact, expressions, and movement from a single reference image.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p>[Youtube] <a href="https://www.youtube.com/watch?v=gg0worsc7bY">Building the AWS of AI Work: HappyRobot CEO Pablo Palafox on Deploying AI Agents Across Enterprise</a> (Bluejay). In this episode, Pablo traces HappyRobot&#8217;s origin story and how HappyRobot went from a handful of people labeling data sets at the kitchen table to running half a million to a million agent runs a day for enterprises like DHL.</p></li><li><p><a href="https://x.com/kwindla/status/2056959360837030344">Gemini 3.5 Flash is too slow for voice agents</a>. Kwindla ran benchmarks on the new Gemini 3.5 Flash and found all Gemini 3 models too slow for real-time voice. Gemini 2.5 Flash is still the production sweet spot.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.11">v1.5.11</a>-<a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.12">v1.5.12</a>. Gemini 3.5 flash, Perplexity Agent API, GPT realtime whisper and several fixes and improvements.</p></li><li><p>Pipecat: No releases</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.66">0.1.66</a>. xAI ASR and TTS, and few small fixes and updates.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - May 18th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-may-18th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-may-18th-2026</guid><pubDate>Mon, 18 May 2026 16:59:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://www.convergence-now.com/enterprise/openai-acquires-weights-gg-ai-voice-cloning-startup/">OpenAI acquires Weights.gg</a>. OpenAI quietly acquihired the team behind Weights.gg (the Replay voice-cloning app, ~$4M raised), which shut down in March 2026. The staff was folded into various OpenAI teams; no standalone product is planned.</p></li><li><p><a href="https://vapi.ai/blog/series-b">VAPI raises $50M Series B</a>. The platform has handled over 1 billion AI voice calls and grown enterprise revenue 10x year-over-year &#8212; with 1M developers building on it and minimal marketing spend behind those numbers.</p></li><li><p><a href="https://venturebeat.com/technology/thinking-machines-shows-off-preview-of-near-realtime-ai-voice-and-video-conversation-with-new-interaction-models">Thinking Machines preview new interaction model</a>. Mira Murati and John Schulman&#8217;s startup previewed TML-Interaction-Small built for full-duplex voice and video conversation. Turn-taking latency hits 0.40 seconds &#8212; natural conversation speed. It scores 77.8 on FD-bench V1.5 vs. 54.3 for Gemini 3.1 Flash Live and 46.8 for GPT-realtime-2.0.</p></li><li><p><a href="https://www.gate.com/news/detail/reactor-releases-a-real-time-3d-world-model-trial-entry-point-with-the-cto-21003004">Reactor Inc releases beta of their real-time AI world generation infrastructure</a>. Ex-Apple and Luma AI founders launched a public beta of real-time AI world model generation &#8212; users explore dynamically rendered 3D environments live in a browser. The CTO demo hit 7.8M views. The company is positioning itself as an infrastructure layer.</p></li><li><p><a href="https://www.pymnts.com/artificial-intelligence-2/2026/better-coms-ai-agent-resolved-35-of-mortgage-calls-alone/">Better.com&#8217;s voice AI agent resolved 35.5% of them without human involvement</a>. Loan officers saved 1,666 hours per month; origination costs dropped 41%; lead-to-lock conversion doubled. Built on ElevenLabs Agents for TTS with lower latency and compliance controls required in mortgage lending.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://deepgram.com/learn/your-restaurant-needs-to-speak-spanish">Deepgram announces Flux Multilingual for Restaurants</a>. Voice-native foundation models and workflows purpose-built for noisy, fast-paced restaurants.</p></li><li><p><a href="https://inworld.ai/speech-to-text">Inworld Realtime STT adds support for Voice Profiling</a>. Inworld&#8217;s STT API now returns a full voice profile alongside the transcript in a single response: emotion (8 categories), vocal style (7 categories), accent, age group, and pitch &#8212; each with a confidence score.</p></li><li><p><a href="https://x.com/DeepgramAI/status/2055000466082177467">Deepgram improves asia-pacific STT support</a>. Nova-3 adds Thai, Cantonese (Traditional), Mandarin (Simplified and Traditional), and Gujarati. Accuracy improvements land on Bengali, Marathi, Tamil, and Telugu at the same time.</p></li><li><p><a href="https://gradium.ai/blog/coval-tts-benchmarks-may-2026">Gradium AI #1 on Coval TTS Benchmarks</a>. With a 158ms median time-to-first-audio and a 2ms interquartile range &#8212; the tightest latency distribution in the field. Word error rate at 3.7%. Cartesia Sonic-3 and ElevenLabs Turbo v2.5 follow.</p></li><li><p><a href="https://livekit.com/blog/answering-machine-detection">LiveKit releases Answering Machine Detection</a>. The feature classifies outbound calls as human, voicemail, IVR, or unavailable within the first second.</p></li><li><p><a href="https://livekit.com/blog/langchain-to-livekit">LiveKit LangChain plugin</a>. The plugin maps LangChain&#8217;s agent abstraction to LiveKit&#8217;s Agents framework, so teams get voice channels on existing implementations.</p></li><li><p><a href="https://github.com/ServiceNow/eva">ServiceNow releases EVA-Bench</a>. An evaluation framework specifically for enterprise voice agents. It surfaces failure modes that generic benchmarks miss.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://developer.nvidia.com/blog/how-to-build-in-vehicle-ai-agents-with-nvidia-from-cloud-to-car/">How to Build In-Vehicle AI Agents</a> (NVIDIA). NVIDIA&#8217;s engineering blog walks through three hardware architectures for cabin AI. Each supports 7B+ parameter models with sub-500ms latency. Covers the full stack from NeMo training to TensorRT-LLM edge deployment.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.9">v1.5.9</a>-<a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.10">v1.5.10</a>. Introducing Answering Machine Detection, Rime Coda TTS, Speechmatics support in Inference, Perplexity LLM and tons of fixes and small improvements.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.2.0">v1.2.0</a>-<a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.2.1">v1.2.1</a>. Smarter turn completion + incomplete turn filtering improvements, OpenAI Realtime reasoning support, and a very long list of fixes and improvements.</p></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - May 11th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-may-11th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-may-11th-2026</guid><pubDate>Mon, 11 May 2026 20:37:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://techcrunch.com/2026/05/05/elevenlabs-lists-blackrock-jamie-foxx-and-longoria-as-new-investors/">ElevenLabs surpassed $500M in AAR</a>. Series D expands past $550M as BlackRock, NVIDIA, Jamie Foxx, and 30+ other investors join. ARR crossed $500M in Q1 2026, up from $350M at year-end 2025 &#8212; $100M net-new ARR in a single quarter. Valuation now at $11B.</p></li><li><p><a href="https://www.linkedin.com/posts/retellai_retell-crossed-60m-arr-and-is-now-emerging-activity-7458171277539717120-3jx2/">Retell AI passed $60M in AAR</a>. 60M ARR with a 35-person team, processing more calls per second than the U.S. 911 system making it one of the fastest trajectories in voice AI infrastructure.</p></li><li><p><a href="https://techcrunch.com/2026/05/06/ethos-raises-22-75m-from-a16z-for-its-expert-network-with-voice-onboarding/">Ethos raises $33M for an expert network onboarding through Voice AI interviews</a>. a16z leads a $22.75M Series A for an expert network that onboards 35,000 people per week through voice AI interviews. The company is on track for eight-figure ARR, charging 30%+ per project to hedge funds, PE firms, and AI labs.</p></li><li><p><a href="https://www.prnewswire.com/news-releases/greenhouse-has-entered-into-a-definitive-agreement-to-acquire-ezra-ai-labs-bringing-conversational-ai-to-the-hiring-process-302762658.html">Greenhouse acquires Ezra AI Labs to embed voice AI interviewing into its ATS</a>. The trigger: applications per recruiter on Greenhouse have spiked 412% since 2023. Ezra generates structured, role-specific interview scores and transcripts with full explainability.</p></li><li><p><a href="https://elevenlabs.io/blog/mahindra">Mahindra deploys voice agents with ElevenLabs to scale outreach for SUV launch</a>. Mahindra deployed ElevenLabs voice agents for the XUV 7XO launch to manage peak demand &#8212; achieving higher contact rates and ~8% conversion uplift during the campaign.</p></li><li><p><a href="https://techcrunch.com/2026/05/09/voice-ai-in-india-is-hard-wispr-flow-is-betting-on-it-anyway/">Wispr Flow bets India is its fastest-growing market</a>. The voice dictation app bets India is its fastest-growing market, adding Hinglish support and hitting 2.5M global downloads, with India growing at 100% month-over-month. India is 14% of installs but only 2% of revenue &#8212; the monetization gap is the real story.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/">OpenAI releases three Voice models</a>. GPT-Realtime-2 with GPT-5-class reasoning and a 128K context window, GPT-Realtime-Translate for live speech translation across 70+ input languages into 13 output languages, and GPT-Realtime-Whisper for streaming STT. The Realtime API exits beta and goes generally available.</p></li><li><p><a href="https://inworld.ai/blog/realtime-tts-2">Inworld releases Realtime TTS-2</a>. A closed-loop voice model that takes prior audio turns as input &#8212; not just transcripts &#8212; to read the user&#8217;s actual tone and pacing, then adapts delivery mid-conversation. Voice Direction lets developers steer output in plain English. Sub-200ms first-chunk latency, 100+ languages, ranked #1 on Artificial Analysis Speech Arena.</p></li><li><p><a href="https://kyutai.org/blog/2026-05-04-pocket-tts-multilingual">Pocket TTS now supports six languages</a>. Kyutai&#8217;s 100M-parameter open-source TTS model goes multilingual: French, Spanish, Portuguese, Italian, German, plus an improved English model &#8212; all running real-time on CPU without a GPU. Single open-source release.</p></li><li><p><a href="https://krisp.ai/blog/viva-2-0-ai-infrastructure-for-voice-ai-agents/">Krisp introduces VIVA 2.0</a>. Turn Prediction v3 detects 47% more true turn-shifts within 200ms vs. v2 across 12+ languages. New Interrupt Prediction model distinguishes real interruptions from backchannel feedback (&#8221;yeah&#8221;, &#8220;okay&#8221;) with under 6% false positives. CPU-only, no transcription required, ~15ms added latency.</p></li><li><p><a href="https://www.twilio.com/en-us/blog/products/signal-2026-product-announcements">Twilio launches a new Conversation Layer</a>. At SIGNAL 2026, Twilio launched Conversation Memory, Orchestrator, and Intelligence &#8212; plus open-source Agent Connect for plugging any AI provider into Twilio voice and messaging channels. Memory builds a living, identity-resolved customer profile that persists across voice, SMS, WhatsApp, and chat.<br></p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://vapi.ai/blog/voice-agent-playbook">Vapi Voice Agent Playbook</a>. 32 chapters distilled from 300M+ calls. Covers voice agent design, deployment, and scaling to production. Practical, opinionated, and grounded in real call data.32 chapters distilled from 300M+ calls. Covers voice agent design, deployment, and scaling to production. Practical, opinionated, and grounded in real call data.</p></li><li><p><a href="https://openai.com/index/delivering-low-latency-voice-ai-at-scale/">OpenAI WebRTC Infrastructure Playbook</a>. OpenAI published its WebRTC architecture details: a split relay + transceiver design handling 900M+ weekly active users at 300&#8211;500ms latency. The relay layer is stateless; the transceiver service owns stateful ICE and DTLS sessions.</p></li><li><p><a href="https://moq.dev/blog/webrtc-is-the-problem/">OpenAI&#8217;s WebRTC Problem</a>. Provocative post in response from the previous one from OpenAI arguing WebRTC is the wrong protocol for voice AI: it aggressively drops audio packets to minimize latency, the opposite of what you want when a 200ms wait beats a dropped prompt.</p></li><li><p>[YouTube] <a href="https://www.youtube.com/watch?v=hbzi15PYzI0">AssemblyAI CEO Dylan Fox on Skywatch</a>. Dylan Fox discusses background noise handling, speaker identification, and what he calls the &#8220;intelligent listening layer&#8221; &#8212; understanding not just what was said but how and in what context. Useful for anyone thinking about the full audio intelligence stack.</p></li><li><p><a href="https://dev.to/vinodsrajpurohit/tts-models-for-indian-languages-the-tech-giving-bharat-a-voice-1ij7">TTS Models for Indian Languages: The Tech Giving Bharat a Voice</a>. Developer survey covering Hindi, Tamil, Bengali, and Telugu TTS models with architecture comparisons and demo links. Good reference for anyone building voice AI for South Asian markets.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.8">v1.5.8</a>. Soniox TTS, new Inworld TTS module, many fixes and upgrades.</p></li><li><p>Pipecat: No releases.</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.64">v0.11.64</a> - <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.65">v0.11.65</a>. Deepgram TTS and some fixes and upgrades.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - May 4th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-may-4th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-may-4th-2026</guid><pubDate>Mon, 04 May 2026 16:55:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://www.microsoft.com/en-us/microsoft-copilot/blog/copilot-studio/extend-ai-voice-support-introducing-real-time-voice-agents-in-microsoft-copilot-studio/">Introducing real-time voice agents in Microsoft Copilot Studio</a>. General availability of real&#8209;time voice agents in Microsoft Copilot Studio launching in Dynamics 365 Contact Center.</p></li><li><p><a href="https://www.nojitter.com/contact-centers/amazon-connect-becomes-connect-customer-">Amazon Connect becomes &#8216;Connect Customer&#8217;</a>. AWS announced Amazon Connect is expanding into four agentic AI solutions: Decisions, Talent, Customer and Health.</p></li><li><p><a href="https://www.alizila.com/leading-chinese-automakers-announce-integration-with-qwen-at-2026-beijing-auto-show/">Alibaba Qwen in Chinese vehicles</a>. Nine automakers announced Qwen integration at the Beijing Auto Show. Drivers can book hotels, order food, and track parcels by voice using Qwen-Omni running on an edge+cloud architecture.</p></li><li><p><a href="https://aithority.com/machine-learning/tells-launches-ai-voice-agents-on-existing-sms-numbers-with-one-click/">Tells turns the same number used for SMS into a natural AI voice agent in one click</a>. Tells.co launched AI Voice Agents, a new capability that lets a business activate a real, natural voice agent on the exact same phone number it already uses for SMS. Activation is a single toggle in the Tells dashboard.</p></li><li><p><a href="https://www.tomsguide.com/ai/ai-voice-cloning-is-everywhere-heres-why-taylor-swifts-new-legal-shield-is-a-blueprint-for-your-digital-safety">Taylor Swift files sound trademark to protect against voice cloning</a>. First major celebrity to use sound marks specifically as an AI cloning defense.</p></li><li><p><a href="https://aiola.ai/blog/field-sales-ai-real-world-challenges-solutions/">AI in Field Sales: Real World Challenges and Solutions from aiOla</a>. Walkthrough of why standard voice AI fails for mobile sales reps: noisy environments, no stable connection, zero desk time.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://deepgram.com/learn/introducing-flux-multilingual">Introducing Flux Multilingual: One Conversational Speech Model for Global Voice Agents</a>. Deepgram&#8217;s first multilingual real-time STT: 10 languages in a single API endpoint. Native turn detection and code-switching, streaming latency under 400ms.</p></li><li><p><a href="https://blogs.nvidia.com/blog/nemotron-3-nano-omni-multimodal-ai-agents/">NVIDIA Launches Nemotron 3 Nano Omni Model</a>. Open multimodal model that handles video, audio, image, and text natively &#8212; hearing tone and background noise rather than reading a transcript. Claims 9x higher throughput than comparable open omni models.</p></li><li><p><a href="https://www.assemblyai.com/products/voice-agent-api">AssemblyAI Releases Voice Agent API</a>. Unified WebSocket pipeline covering STT (Universal-3 Pro Streaming), LLM reasoning, TTS, turn detection, and interruption handling in one connection.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://poly.ai/blog/barge-in-voice-ai-interruption-handling">How barge-in handling impacts the quality of your voice AI</a> (Poly.ai). PolyAI&#8217;s deep-dive on barge-in handling: interruptions happen in roughly 1 in 5 calls, and false positives &#8212; triggered by background noise &#8212; do more damage to caller trust than missed ones.</p></li><li><p><a href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/voice_ai_agents/insurance_claim_live_agent_team">Insurance Claim Live Agent Team example</a> (Awesome LLM Apps). Good example from Shubham Saboo on how to use ADK and Gemini Live for extracting structured information from a live conversation.</p></li><li><p>[YouTube] <a href="https://www.youtube.com/watch?v=NvaLzcSvYyQ">Pipecat 1.0</a> (Pipecat TV). The Pipecat core team celebrates the release of Pipecat 1.0, a huge milestone after two years and 100+ releases. The crew dives into their favorite features, what made it into 1.0, and some of the challenges along the way. </p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.7">v1.5.7</a>. SLNG support, Timed transcriptions, LLM gateway priorities, gpt-5.4-mini,  and many fixes, configuration options and upgrades.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.1.0">v1.1.0</a>. New STT providers (Mistral Voxtral, xAI) + streaming TTS (xAI, Soniox) for lower-latency voice agents, Deegram Flux multi-language, faster turn taking and many fixes and updates.</p></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Apr 27th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-apr-27th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-apr-27th-2026</guid><pubDate>Tue, 28 Apr 2026 18:19:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://www.tomsguide.com/ai/google-meets-ai-note-taking-feature-can-now-summarize-your-in-person-meetings-heres-how-it-works">Google Meet AI note-taking for in-person meetings</a>. Gemini&#8217;s &#8220;Take Notes for Me&#8221; now captures face-to-face meetings, not just video calls &#8212; tap the button on Android, set the phone near the room, and get a transcript, summary, and action items saved to Drive.</p></li><li><p><a href="https://venturebeat.com/business/synthflow-ai-and-8x8-enter-strategic-partnership-to-deliver-next-generation-agentic-ai">Synthflow AI partnership with 8x8</a>. Synthflow has processed 65M+ voice interactions to date; 8x8 plans to resell it through channel partners and its App Store for SMBs.</p></li><li><p><a href="https://investors.soundhound.com/news-releases/news-release-details/soundhound-ai-acquire-liveperson-combining-proprietary-voice">SoundHound AI To Acquire LivePerson</a>. Combining Proprietary Voice Agentic AI and Digital Messaging to Create a World Leading End-to-End Omnichannel Conversational AI Platform.</p></li><li><p><a href="https://www.ericsson.com/en/blog/2026/4/ai-voice-in-telecom-powering-calls-and-securing-networks">Ericsson is integrating AI capabilities into IMS voice calling</a>. Ericsson Ventures&#8217; investments in Cartesia and Hiya highlight the network&#8217;s dual role in the AI voice era: enabling ultra&#8209;low&#8209;latency AI calling while safeguarding trust through native fraud protection.</p></li><li><p><a href="https://www.cnbc.com/2026/04/21/volkswagen-voice-ai-chinese-cars-automaker.html">Volkswagen Voice AI in China</a>.  All China-built VW models shipping H2 2026 will run an on-device LLM-powered voice agent drawing on tech from Tencent, Alibaba, and Baidu. The model runs entirely on the car, not the cloud. Announced at Auto China 2026 alongside a broader agentic AI roadmap for the market.</p></li><li><p><a href="https://www.tvtechnology.com/production/adobe-and-speechmatics-deliver-cloud-grade-on-device-speech-recognition-for-premiere">Adobe and Speechmatics Deliver `Cloud-Grade&#8217; On-Device Speech Recognition for Premiere</a>. Companies handling content before it goes public can now work seamlessly from anywhere: on a film set, between client meetings, on a flight&#8212;at full accuracy, with no dependency on a connection and no interruption to the work.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://x.ai/news/grok-voice-think-fast-1">xAI releases Grok Voice Think Fast 1.0</a>. Debuts at #1 on the &#964;-voice Bench with 67.3%, more than 20 points ahead of Gemini Flash Live (43.8%) and GPT Realtime (35.3%). Already live as Starlink&#8217;s phone agent.</p></li><li><p><a href="https://www.gizmochina.com/2026/04/24/xiaomi-introduces-mimo-v2-5-tts-and-asr-as-a-full-voice-pipeline-for-the-agent-era/">Xiaomi introduces MiMo-V2.5-ASR</a>.  Xiaomi completes the full voice pipeline with open-source ASR covering Mandarin, English, Chinese dialects, and code-switched speech. ASR weights available on Hugging Face.</p></li><li><p><a href="https://elevenlabs.io/blog/introducing-agent-templates">ElevenLabs introduces Agent Templates</a>. Templates map to the use cases where agents drive the most value: customer support, sales, operations, and internal enablement.</p></li><li><p><a href="https://x.com/usebland/status/2047768691564232795">Bland users get a free phone number on signup now</a>. </p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://livekit.com/blog/gemini-3.1-flash-tts-prompting-guide">A Practical Guide to Prompting Gemini 3.1 Flash TTS</a>. A working set of rules for getting natural, emotionally appropriate speech out of gemini-3.1-flash-tts-preview.</p></li><li><p><a href="https://medium.com/@kauxhik77/meet-mimi-the-neural-audio-codec-powering-the-next-generation-of-speech-llms-22a77f3b9a38">Mimi Codec deep-dive</a>. Walkthrough of Kyutai&#8217;s neural audio codec behind Moshi and Sesame CSM: a split RVQ that isolates semantic tokens in one quantizer while seven acoustic quantizers handle fidelity. Fully causal, 80ms latency, 1.1 kbps at 12.5 Hz. Good primer on audio tokenization if you&#8217;re building or evaluating speech LLMs.</p></li><li><p><a href="https://www.technology.org/2026/04/24/streaming-tts-models-fail-over-60-of-sentences-containing-numbers-dates-and-prices/">Streaming TTS Models Fail Over 60% of Sentences Containing Numbers, Dates, and Prices</a>. Root cause analysis: streaming mode gives models 5&#8211;20x less context than batch, forcing pronunciation decisions before a full sentence arrives. Failure rates spike on phone numbers, alphanumeric IDs, and addresses. Practical takeaway for voice agent builders handling financial or contact-center content.</p></li><li><p><a href="https://techcrunch.com/2026/04/24/nothing-introduces-an-ai-powered-dictation-tool/">How AstroBeam built the world&#8217;s first voice-only VR game</a>. AstroBeam is on a mission to make voice a first-class input in games. Their debut title, Stellar Cafe, is available on Meta Quest and is fully playable using only your voice.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.5">v1.5.5</a> -  <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.6">v1.5.6</a>. xAI TTS, Deepgram Flux multilanguage, Inworld STT, Pulse STT, Simplismart Qwen 3 TTS and some fixes and upgrades.</p></li><li><p>Pipecat: No releases.</p><ul><li><p>Pipecat Cloud: New generic websocket interface for server-server use cases.</p></li></ul></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Apr 20th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-apr-20th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-apr-20th-2026</guid><pubDate>Mon, 20 Apr 2026 18:23:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://elevenlabs.io/blog/razorpay">Razorpay scales outbound merchant engagement with ElevenAgents</a>. India&#8217;s omnichannel payments platform for businesses, deployed ElevenAgents to power outbound voice agents in Hinglish achieving close to 28% connection rate - on par with human call center benchmarks.</p></li><li><p><a href="https://www.speechmatics.com/company/articles-and-news/ai-can-now-understand-health-signals-from-15-seconds-of-your-voice-including-fatigue-stress-and-type-2-diabetes">Thymia partners with Speechmatics to provide voice biomarker intelligence</a>. Their joint platform identifies 30-plus health signals from 15 seconds of natural speech, including stress, fatigue, depression and anxiety symptoms, type 2 diabetes and driver impairment, then acts on them in real time.</p></li><li><p><a href="https://www.8x8.com/products/ai-studio?utm_medium=social-media-organic&amp;utm_source=linkedin">8x8 releases 8x8 AI Studio</a>. With 8x8 AI Studio, you describe what you need and the Builder creates, tests, and deploys it. From your first conversational agent to fully autonomous workflows that take action across your business systems end to end.</p></li><li><p><a href="https://techcrunch.com/2026/04/16/deepl-known-for-text-translation-now-wants-to-translate-your-voice/">DeepL, known for text translation, now wants to translate your voice</a>. DeepL, a translation company , released a voice-to-voice translation suite that covers use cases like meetings, mobile and web conversations, and group conversations for frontline workers through custom apps.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://www.indianweb2.com/2026/04/weya-ai-launches-hush-lightweight-open.html">Weya AI Launches &#8216;Hush&#8217;: Lightweight Open-Source Speech Enhancement Model for Voice AI</a>.  An open-source, real-time speech enhancement model for Voice AI. It&#8217;s unique because it suppresses background speakers, even at 8 MB and running efficiently on a CPU.</p></li><li><p><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts/">Gemini 3.1 Flash TTS Preview</a>.  The Gemini 3.1 Flash TTS Preview model introduces expressive audio tags for controlling narration, as well as overall improvements to naturalness, controllability, and multilinguality.</p></li><li><p><a href="https://3d-models.hunyuan.tencent.com/world/">Tencent releases HY World 2.0 3d model</a>. Input a text description or an image, and the model synthesizes high-fidelity, roamable 3D Gaussian Splattings (3DGS) / Mesh scenes. Generated worlds support free navigation with physical collision, seamlessly compatible with Unity/UE engines.</p></li><li><p><a href="https://x.com/ModelScope2022/status/2043605089441489263?s=20">MOSS-TTS-Nano MOSI.AI and OpenMOSS</a>. Designed for realtime speech generation without a GPU. Runs directly on CPU, keeping the deployment stack simple enough for local demos, web serving, and lightweight product integration.</p></li><li><p><a href="https://x.com/retellai/status/2044817955398287668">Retell AI launches pre-approved SMS numbers</a>. Your agent can now send SMS during a call without going through A2P verification.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://arxiv.org/abs/2602.12249">&#8220;Sorry, I Didn&#8217;t Catch That&#8221;: How Speech Models Miss What Matters Most </a>(Together AI) 15 models from OpenAI, Deepgram, Google, and Microsoft were tested on street name transcription from linguistically diverse US speakers.</p></li><li><p>[Youtube] <a href="https://www.youtube.com/watch?v=EZro3CcGALA">Built to Ship with Kwindla - CoFounder @ Daily.co</a> (Voker) Voker CEO Tayler talks with Kwindla about the path from WebRTC to pipecat the future of Voice AI.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.3">v1.5.3</a> -  <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.4">v1.5.4</a>. Krisp Viva SDK, Runway avatar plugin, Cerebras LLM plugin, xAI STT, gemini-3.1-flash-tts-preview, xAI Grok LLM for inference and many other fixes. </p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v1.0.0">v1.0.0</a>. Pipecat 1.0 includes a bunch of new additions, bug fixes, and deprecation removals.</p><ul><li><p>Pipecat subagents 0.1.0: New  distributed multi-agent framework for Pipecat. Each agent runs its own pipeline and communicates through a shared message bus. This lets you decompose complex systems into specialized agents that can run locally or across machines.</p></li><li><p>Pipecat Flows 1.0.0: New version updated for Pipecat 1.0 with many breaking changes and dependency upgrades.</p></li></ul></li><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Apr 13th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-apr-13th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-apr-13th-2026</guid><pubDate>Mon, 13 Apr 2026 14:12:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://elevenlabs.io/on-prem-deployments">ElevenLabs introduces On-Premise and On-Device</a>. ElevenLabs can now be deployed on-premise and on-device. This expands our deployment options beyond cloud and VPC, to cover the full range of enterprise environments.</p></li><li><p><a href="https://seed.bytedance.com/en/blog/introducing-seed-full-duplex-speech-llm-attentive-listening-robust-interference-suppression-enabling-more-natural-interaction">ByteDance introduced Seeduplex</a>. A native full-duplex end-to-end speech LLM with attentive listening, robust interference suppression enabling more natural interaction.</p></li><li><p><a href="https://junxuan-li.github.io/lca/">Meta presents Large-Scale Codec Avatars (LCA)</a>. A high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner. </p></li><li><p><a href="https://mistral.ai/news/voxtral-tts">Mistral open sourced Voxtral TTS</a>.  Mistral open sourced Voxtral TTS, a text-to-speech model that clones voices from 3 seconds of audio and runs on edge devices.</p></li><li><p><a href="https://github.com/livekit/livekit-wakeword">LiveKit releases WakeWord library</a>. An open-source wake word library for creating voice-enabled applications.</p></li><li><p><a href="https://www.bland.ai/blogs/fluent-next-generation-multilingual-transcription-voice-agents">Bland introduces</a><strong><a href="https://www.bland.ai/blogs/fluent-next-generation-multilingual-transcription-voice-agents"> </a></strong><a href="https://www.bland.ai/blogs/fluent-next-generation-multilingual-transcription-voice-agents">Fluent STT</a>. Next-Generation Multilingual Transcription for Voice Agents providing 27% reduction in transcription error.</p></li><li><p><a href="https://daniellin94144.github.io/FDB-v3-demo/">Full-Duplex-Bench-v3 (FDB-v3) released</a>.  A benchmark combining real human disfluent speech with multi-step tool use to evaluate voice agents under realistic conditions.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://medium.com/@ggarciabernardo/video-ai-avatars-why-we-need-to-ditch-the-meeting-room-architecture-1201249d807d">Video AI Avatars: Why We Need to Ditch the Meeting Room Architecture</a> (LiveTok) Existing platforms are designed for video conferencing between humans and not for 1:1 conversation with an Agent.   This post explores the implications and limitations of the current approaches.</p></li><li><p>[Youtube] <a href="https://www.youtube.com/watch?v=mIm1jIezSGU">From Dark Matter to Deep Learning: Deepgram CEO Scott Stephenson on the Future of Voice AI</a> (BlueJay) Faraz and Scott dig into why the "bigger is better" model narrative is fundamentally flawed, why the real battle in AI isn't small vs. large models but real-time vs. not, and how the laws of physics themselves will shape the architecture of voice AI systems at scale.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.2">v1.5.2</a>. One of the versions with more changes I've seen.   Lot of fixes and improvements and new capabilities to many providers.  Also support for Gemini 3.1 Flash Live, Voxtral TTS and answering machine detection. </p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v0.0.108">v0.0.108</a>. Added new services like Sarvam, Novita, xAI TTS, and Smallest AI TTS.  Better handling of interruptions, race conditions, and context edge cases and many bugfixes.</p></li></ul><ul><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Mar 23rd 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-mar-23rd-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-mar-23rd-2026</guid><pubDate>Mon, 23 Mar 2026 18:08:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://decagon.ai/blog/introducing-duet">Decagon launches Duet</a>. Duet automatically generates test scenarios, simulates conversations, stresses edge cases, and flags situations where the agent may fail. Also investigates live conversations and tunes instructions accordingly.</p></li><li><p><a href="https://agoras-website.webflow.io/en/extensions/palabra-ai">Palabra.ai brings real-time multilingual communication to any Agora-powered experience</a>. Developers can add Palabra as a drop-in extension to enable real-time multilingual voice conversations in their Agora-powered apps - from live streaming and social platforms to education and customer support.</p></li><li><p><a href="https://docs.roark.ai/documentation/recipes/accent-detection">Roark (YC W25) releases accent detection for voice AI calls</a>. A dedicated ML model trained on 500k+ labeled speech samples, reaching ~89% accuracy across 15 English accent variants.</p></li><li><p><a href="https://www.heidihealth.com/en-us/blog/how-heidi-improved-asr-nvidia-nemotron">How Heidi cut ASR costs 64% and latency 75% with NVIDIA Nemotron Open ASR</a>. The Clinical AI platform moved beyond "off-the-shelf" APIs, achieving a level of performance and efficiency that was previously considered unattainable at this scale.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://artificialanalysis.ai/articles/nemotron-3-voicechat-leader-speech-pareto">NVIDIA early access to Nemotron 3 VoiceChat</a>. A ~12B parameter Speech to Speech model that leads the open weights Conversational Dynamics vs. Speech Reasoning pareto frontier.</p></li><li><p><a href="https://livekit.com/blog/adaptive-interruption-handling">LiveKit releases Advanced Interruption Handling</a>.  A new model to discriminate between genuine attempts to interrupt and incidental speech or noise.</p></li><li><p><a href="https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-bedrock-webrtc/">Amazon Bedrock AgentCore Runtime adds WebRTC support</a>. AgentCore now supports WebRTC for real-time bidirectional streaming between clients and agents, adding to the existing WebSocket support. With WebRTC, developers can improve latency and user experience on non-ideal networks.</p></li><li><p><a href="https://herimor.github.io/voxtream2/">VoXtream2: Full-stream TTS with dynamic speaking rate control</a>. A zero-shot full-stream TTS model with dynamic speaking-rate control that can be updated mid-utterance on the fly.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://www.ultravox.ai/voice-ai/voice-ai-glossary-key-terms-and-concepts-explained">Voice AI Glossary: Key Terms and Concepts Explained</a> (Ultravox). Ultravox created this list of key terms and concepts you're likely to encounter if you're building voice agents.</p></li><li><p><a href="https://blog.cartesia.ai/p/mamba-3">Mamba-3: An Inference-First State Space Model</a>. SSMs marked a major advance for the efficiency of modern LLMs. Mamba-3 takes the next step, shaping SSMs for a world where AI workloads are increasingly dominated by inference.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.5.0">v1.5.0</a>. Major turn handling improvements (adaptive interruption handling and dynamic endpointing), smarter context summarization, more metrics and a new plugin to detect blocked threads.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v0.0.106">v0.0.106</a>. Wake phrase turn detection, improved Perplexity LLM support and Daily DTMFs support, better performance, stability &amp; reliability fixes.</p></li></ul><ul><li><p>TEN Framework: No releases.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Mar 16th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-mar-16th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-mar-16th-2026</guid><pubDate>Mon, 16 Mar 2026 16:41:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://investor.agora.io/news-releases/news-release-details/agora-removes-barriers-scalable-voice-ai-agents">Agora releases ConvoAI Studio</a>. A no-code deployment tool to configure, test, &amp; deploy Voice AI agents.  In addition Agora has released more targeted Voice AI Agent solutions for Customer Support and Sales &amp; Marketing Agents.</p></li><li><p><a href="https://www.hume.ai/blog/opensource-tada">Hume AI releases its first open source TTS model, TADA</a>. The model is based on a novel tokenization schema that synchronizes text and speech one-to-one to get the fastest LLM-based TTS system available, with competitive voice quality, virtually zero content hallucinations, and a footprint light enough for on-device deployment.</p></li><li><p><a href="https://fish.audio/s2/">Fish Audio S2 TTS launched</a>.  Open-weights voice AI model with the highest level of control and emotions.</p></li><li><p><a href="https://www.together.ai/blog/build-real-time-voice-agents-on-together-ai">Together AI announced a full suite of capabilities for building real-time voice agents</a>.  Co-located STT, LLM, and TTS on one cloud, eliminating inter-vendor network hops for end-to-end pipeline latency under 700ms. Cartesia Sonic-3 (TTS) and Deepgram (STT) are now natively hosted.</p></li><li><p><a href="https://chatgpt.com/apps/retell-ai/asdk_app_695b101534508191a313998a4a5badc0">Retell AI is now available directly in ChatGPT</a>. The application simplifies building, deploying and managing AI voice agents directly from ChatGPT.</p></li><li><p><a href="https://github.com/dograh-hq/dograh">Dograh Voice AI Platform released</a>.  Opensource platform similar to VAPI or RetellAI and built on top of Pipecat.</p></li><li><p><a href="https://github.com/huggingface/speech-to-speech">Hugging Face&#8217;s speech-to-speech is back from the dead</a>. Open and modular pipeline from Hugging Face is active again to run local  voice agents by mixing and matching the best open-source models.   Just added 5 modern STT/TTS backends: Qwen3-TTS, Parakeet, Kokoro, Pocket TTS, and MLX Whisper. </p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://www.daily.co/blog/nvidia-nemotron-3-super/">NVIDIA Nemotron 3 Super, a new open source model for voice AI and agentic tasks</a> (Daily). The new open source LLM launched by NVIDIA performs at the same level as the new GPT-5.4 model and developers now have a meaningful complete open stack for realtime voice from NVIDIA with Nemotron 3 Nano, Nemotron Speech ASR, Nemotron 3 Super. Open models, open training data.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: <a href="https://github.com/livekit/agents/releases/tag/livekit-agents%401.4.5">v1.4.5</a>.  GPT-5.4, Phonic model, NVIDIA STT diarization, support to split instructions per modality and tons of fixes and small improvements.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v0.0.105">v0.0.105</a>. Better service settings, concurrent audio contexts, automatic service failover, improved system instructions support for sharing LLM context across services, and many improvements and fixes.</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.60">0.11.60</a> - <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.62">0.11.62</a>. Deepgram flux STT support, small fixes and chores.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Mar 9th 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-mar-9th-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-mar-9th-2026</guid><pubDate>Mon, 09 Mar 2026 17:55:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://www.businesswire.com/news/home/20260305577534/en/Smart-Glasses-Rebuilt-for-Privacy-Brilliant-Labs-Neuphonic-TheStage-AI-Move-AI-Off-the-Cloud">Smart Glasses, Rebuilt for Privacy: Brilliant Labs, Neuphonic &amp; TheStage AI Move AI Off the Cloud</a>. Neuphonic voice models will run on TheStage AI on-device inference engine, for Brilliant Labs glasses.</p></li><li><p><a href="https://www.linkedin.com/posts/davidezhang_openclaw-spectacles-arvr-ugcPost-7436147175761420288-mzJq">Wired Snap Spectacles to OpenClaw to prototype a spatial AI assistant</a>.  Prototype including pinch to talk, on-device speech-to-text, and reply shows in AR. Ask it to open apps, search the web, draft documents, all hands-free on the go.  <a href="https://github.com/davidezhang/openclaw-spectacles">Opensource repo</a>.</p></li><li><p><a href="https://cartesia.ai/customers/lorikeet">Lorikeet&#8217;s AI Concierge Resolves Complex Customer Issues with Cartesia</a>. Lorikeet uses Cartesia voices in its agents to build trust through natural pacing, emotional balance, and carefully tuned disfluencies. Deployed voice agents in &lt;24 hours in a food crisis, handling 21,000+ calls over 9 days with 100% uptime.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p><a href="https://docs.nvidia.com/nim/speech/latest/reference/support-matrix/asr.html#parakeet-1-1b-rnnt-multilingual">NVIDIA Parakeet-RNNT-1.1B-Multilingual model now offers three distinct profile configurations</a>. These options range from a 25-language auto-detection mode to specialized profiles optimized for specific language prompts and Indic dialects.</p></li><li><p><a href="https://www.assemblyai.com/universal-3-pro-streaming">AssemblyAI Universal-3-Pro is now available for streaming</a>. Bringing AssemblyAI&#8217;s most accurate speech model to live audio for the first time.</p></li><li><p><a href="https://x.com/kamath_sutra/status/2028693153629491595">Introducing a new S2S model called Hydra</a>.  A native speech-to-speech model that doesn't wait for turn-taking, doesn't flatten emotion into text, and doesn't break when you interrupt it mid-sentence.  Demo in the link.</p></li><li><p><a href="https://github.com/FireRedTeam/FireRedVAD">FireRedVAD: A SOTA Industrial-Grade Voice Activity Detection &amp; Audio Event Detection</a>.  According to the author outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD.</p></li><li><p>Retell AI voice agents now dynamically match the caller&#8217;s speaking pace. <a href="https://x.com/retellai/status/2030000210336960741?s=20">Here is the demo</a>.</p></li><li><p><a href="https://livekit.io/ui">LiveKit launches Agents UI</a>.  An open-source shadcn component library for building polished React frontends for your voice agents including audio visualizers, media controls, session management tools and chat transcripts.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://www.ntik.me/posts/voice-agent">How I built a sub-500ms latency voice agent from scratch</a> (Nick Tikhonov).  How building a custom a voice agent pipeline using DeepGram Flux and Groq can beat the latency of popular commercial platforms.</p></li><li><p><a href="https://www.youtube.com/watch?v=T45HOvl3ue0">AI Missions, Agent Identity &amp; the Future of Voice Infrastructure</a> (Bluejay).  David Casem, Co-founder and CEO of Telnyx, for a wide-ranging conversation on what it really takes to build telecom infrastructure and the shift from AI workflows to AI missions &#8212; long-horizon, multi-step tasks that agents execute autonomously over hours or days.</p></li><li><p><a href="https://www.youtube.com/watch?v=chXqb-DdeWc">How to Build and Scale Voice Agents Using NVIDIA Nemotron, Modal, and Daily</a> (NVIDIA Nemotron Labs). Walk through how to build and scale real-time voice agents using NVIDIA&#8217;s open Nemotron models, orchestrated on Modal and powered by Daily for real-time audio and agent communication.</p></li><li><p><a href="https://www.ultravox.ai/blog/what-we-need-to-make-voice-ai-fully-agentic">What we need to make voice AI fully agentic</a> (Ultravox). Why agentic use cases are well on their way to dominating the world of text models&#8211;think of Claude Code&#8217;s meteoric rise&#8211;many production voice-based systems remain stuck in late 2024 and what do we need to move forward.</p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: v1.4.4.  Model upgrades to Sonic 3 and GPT-Realtime 1.5, expands the ecosystem with Telnyx, SambaNova, and Keyframe Labs plugins, and enhances performance through optimized WAV decoding and advanced AssemblyAI features.</p></li><li><p>Pipecat: <a href="https://github.com/pipecat-ai/pipecat/releases/tag/v0.0.104">v0.0.104</a>. Enhances observability through detailed latency metrics (TTFB, text aggregation, and startup timing), sophisticated LLM context summarization controls, and expanded support for Azure private endpoints and LemonSlice avatars.</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.56">0.11.56</a> - <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.59">0.11.59</a>. Small fixes and chores.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Updates - Mar 2nd 2026 ]]></title><description><![CDATA[Weekly Voice and Video AI Product and Platform news]]></description><link>https://livetok.substack.com/p/weekly-updates-mar-2nd-2026</link><guid isPermaLink="false">https://livetok.substack.com/p/weekly-updates-mar-2nd-2026</guid><pubDate>Mon, 02 Mar 2026 15:26:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3LT1!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b8d00cb-798c-4ea3-9762-5098f762243a_380x380.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>&#128478;&#65039; Market and Product News</strong></h2><ul><li><p><a href="https://www.prnewswire.com/news-releases/elevenlabs-partners-with-google-cloud-for-cloud-services-and-the-latest-nvidia-blackwell-gpus-302697853.html">ElevenLabs Partners with Google Cloud for Cloud Services and the Latest NVIDIA Blackwell GPUs</a>.  Announced a multi-year extension of their strategic collaboration to make high-quality, artificial intelligence voice tools more accessible to businesses worldwide.</p></li><li><p><a href="https://www.cxtoday.com/contact-center/zoom-virtual-agent-3-0-chatbot-resolution/">Zoom Launches Virtual Agent 3.0 to Fix Chatbot&#8217;s Broken Promise</a>.   The virtual agent can now orchestrate multi-step workflows across CRM, billing, order management, and other enterprise systems &#8211; not just respond to a query, but take action across connected tools and see a process through to completion.</p></li><li><p><a href="https://www.perplexity.ai/hub/blog/perplexity-apis-deliver-powerful-ai-to-the-world%E2%80%99s-largest-android-device-maker">Perplexity APIs deliver powerful AI to the world&#8217;s largest Android device maker</a>. Perplexity is now integrated into Samsung Galaxy S26 phones at a deep level, powering search and reasoning for both the Perplexity assistant and Samsung's Bixby.  Galaxy S26 users can say "Hey Plex" to launch the Perplexity assistant.</p></li></ul><ul><li><p><a href="https://thenextweb.com/news/voiceline-raises-e10m-to-scale-its-voice-ai-platform-for-frontline-enterprise-teams">VoiceLine raises &#8364;10M to bring voice-first AI to enterprise frontline teams</a>. VoiceLine, a Munich-based startup that builds voice-first artificial intelligence for enterprise frontline workers, has closed &#8364;10 million in Series A funding to accelerate growth, expand its product and bring its technology to more customers across Europe and beyond.</p></li><li><p><a href="https://newsroom.ibm.com/2026-02-24-deepgram-and-ibm-introduce-advanced-voice-capabilities-for-enterprise-ai">Deepgram and IBM Introduce Advanced Voice Capabilities for Enterprise AI</a>.  To address client needs for highly performant, enterprise-grade transcription and real-time captioning, IBM will embed Deepgram&#8217;s capabilities into watsonx Orchestrate.</p></li></ul><h2><strong>&#129520; Platform News</strong></h2><ul><li><p>Layercode is sunsetting its voice AI platform.  The team will focus on a new product is called <a href="https://77152dd8.click.convertkit-mail2.com/4zu39xzkwesehpvwropaxh6mmrm77a5hl236r/owhkhqhwp88q87cv/aHR0cHM6Ly90b3lvLmFpLw==">Toyo</a>. It's an AI computer in the cloud that helps founders grow their businesses with AI agents, without needing any coding or technical skills.</p></li><li><p><a href="https://developers.openai.com/api/docs/models/gpt-realtime-1.5">gpt-realtime-1.5 is live in Realtime API</a>.   The new model delivers a +5% intelligence lift on Big Bench Audio, which measures reasoning ability, as well as +10.23% on alphanumeric transcription and +7% on instruction following in internal evals.</p></li><li><p><a href="https://rime.ai/resources/rime-slng">Rime is partnering with SLNG to make state-of-the-art voice AI available on demand across underserved markets</a>. Combining Rime&#8217;s latest voice model, Arcana V3, with SLNG&#8217;s global edge infrastructure to deliver low-latency, production-grade voice experiences, deployed locally and built for compliance from day one.</p></li><li><p><a href="https://www.agentconference.com/agenticlist/2026">LiveKit has been recognized as a leading Agent Development Platform for building, testing, and deploying autonomous AI agents on The Agentic List 2026</a>.</p></li><li><p><a href="https://www.linkedin.com/posts/vkhurana2_some-news-wayfaster-is-joining-livekit-activity-7432118628663160834-ypnp/">Wayfaster is joining LiveKit</a>. LiveKit has acquired Wayfaster, the infrastructure that interviewed hundreds of thousands of people looking for their next jobs.</p></li></ul><h2><strong>&#128214; Reading</strong></h2><ul><li><p><a href="https://blog.livekit.io/prompting-voice-agents-to-sound-more-realistic/">Prompting voice agents to sound more realistic</a> (LiveKit). To make a cascaded voice agent sound like a real person on a call your system prompt needs to do two things well: show the model what you mean, and reinforce the same behaviors from multiple angles.</p></li><li><p>[Podcast] <a href="https://voice-ai-newsletter.krisp.ai/p/promptable-speech-language-models?hide_intro_popup=true">Promptable Speech Language Models</a> (Krisp). Dylan Fox (Founder &amp; CEO at AssemblyAI) talks about Universal-3 Pro, the first speech language model optimized specifically for voice AI, goes further with advanced prompting capabilities that let developers customize model behavior for their exact use case.</p></li><li><p><a href="https://blog.livekit.io/voice-agent-skills-for-coding-assistants/">Voice agent skills for coding assistants</a> (LiveKit). Why agent skills for voice agents matter now and how to use LiveKit skill to enhance your developer workflow building voice agents.</p></li><li><p><a href="https://getbluejay.ai/blog/how-testing-evolves-for-full-duplex-(speech-to-speech)-models">How Testing Evolves for Full-Duplex (Speech-to-Speech) Models</a> (BlueJay).   Notes on how testing methodologies are adapting in the world of voice AI and new models where conversations are not modeled as turns anymore. </p></li></ul><h2>&#128230; Releases</h2><ul><li><p>LiveKit Agents: No releases.</p></li><li><p>Pipecat: No releases.</p></li><li><p>TEN Framework: <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.54">0.11.54</a> - <a href="https://github.com/TEN-framework/ten-framework/releases/tag/0.11.55">0.11.55</a>. Whisper STT and improved Sarvam STT, OpenClaw support, Telnyx and Plivo telephony providers and many new examples and fixes.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://livetok.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive next updates</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>