hn.today

The STT-LLM-TTS voice stack is dead

skeptrune.com6 points1 comments
Screenshot of The STT-LLM-TTS voice stack is dead

This describes building call4me, a managed service that lets software agents place real-world phone calls on a user’s behalf. The product requires an MCP URL in the agent and enforces upfront collection of everything a business will ask for (profile fields plus one of 12 category-specific requirement sets) so the caller can only say what it was given. Calls run on Cloudflare Workers with one Durable Object per live call, use Telnyx for PSTN connectivity, and record early traction (about $610 gross and $215 MRR in week one). During calls the system polls for progress, invites the user to join by pressing 1, and posts a transcript-derived recap into the agent context.

The key engineering choice is a speech-to-speech architecture: Telnyx and GPT-Live share the same G.711 μ-law 8 kHz audio, so audio frames are relayed untouched - no STT or TTS - while a separate “back office” text model (gpt-5.5) owns side effects and tools (end_call, press_digits, ask_user, connect_person). Practical work solved phone-menu and routing failures with transcript pattern detection, a MenuRecovery state machine, and sending real DTMF tones as audio for non-US routes. Operational fixes include isolating live calls from deploys, heartbeat/reattach logic, listen-only user join, and periodic filler speech when awaiting user answers. Recommendations: avoid unnecessary audio conversions, split live voice from decision-making, verify model-executed actions, and gather all call data before dialing.

Read on skeptrune.com1 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.