Mutable transcripts are breaking your voice agents
We had a voice agent that worked perfectly in testing. Clean transcripts, fast responses, smooth conversations. Then we put it in production and it started acting like it had memory problems.
The agent would respond to the same question twice, skip parts of multi-step instructions, or suddenly forget what the user just said. We thought it was a context window issue or some weird LLM behavior.
Turns out, it was the speech-to-text.
Most streaming ASR systems work like autocorrect on steroids. As more audio comes in, they revise their earlier guesses. "I need to book a flight" might start as "I need to" then become "I need to book" then "I need to book a light" before finally settling on "I need to book a flight."
This works fine for subtitles. It breaks agents completely.
Here's what was happening: our agent would see "I need to book a light" and start processing that request. Then the transcript would update to "I need to book a flight" and trigger the agent again. Same input, different processing, chaos in the conversation flow.
Classic symptom: your voice agent works great with clear, simple commands but falls apart during natural conversation. The issue isn't intelligence — it's transcript instability.
The fix is append-only transcription. Once text is committed, it never changes. You get stable, monotonic output that agents can actually work with.
We switched to NetEase's Confucius4-R2T2 model, which does true streaming with what they call "Longest Stable Prefix" output. It commits text in chunks and never revises. The latency is around 200-600ms, which is fast enough for conversation, and it's lightweight enough to run locally.
The difference was immediate. Our agent stopped double-processing, stopped losing context mid-conversation, and started feeling responsive instead of confused.
Key implementation details:
- Configure chunk timing based on your use case (we use 400ms for natural conversation)
- Build your agent to expect append-only input — no revision handling needed
- Set up proper audio buffering so you don't lose speech during processing
- Test with overlapping speech and background noise, not just clean audio
The broader lesson: when you're building voice agents, the handoff from speech-to-text to language model is where things break. Mutable transcripts create race conditions that no amount of prompt engineering can fix.
If you're running voice agents in production, audit your ASR pipeline. If it's revising output, you're probably seeing phantom failures that look like agent intelligence problems but are actually transcript coordination problems.