Voice agent troubleshooting starts by locating the failing boundary: telephony, media, ASR, LLM, tools, or TTS. If your voice bot is dropping calls, first establish how long it stayed connected and which system ended it. For audio, response-quality, or latency problems, compare the input and output at each component before changing prompts.
Last Updated: September 25, 2026
For a one-off failure, collect the provider call record, relevant audio, and agent trace before changing anything. For repeated drops, use the elapsed-time signature and evidence packet below. The later checklists cover speech recognition, model behavior, tools, synthesis, network quality, and turn-taking.
TL;DR: Choose the first diagnostic boundary from the observed symptom:
| Decision Signal | First Boundary to Check | Evidence to Capture |
|---|---|---|
| Drops consistently after 20–32 seconds | Initial SIP ACK routing | SIP ladder from both peers, Contact and Record-Route headers |
| Drops at a fixed longer interval | SIP session refresh | Session-Expires, refresher role, re-INVITE or UPDATE, final BYE |
| Drops immediately after sustained silence | Media flow or silence policy | RTP counters, silence event, application timeout, disconnect reason |
| Calls never connect | Provider routing and SIP registration | Provider status events, destination, answer webhook response |
| No sound or garbled audio | Codec, RTP, and media negotiation | Codec selection, packet counters, media events |
| Wrong responses or tool timeouts | Agent runtime and dependencies | Trace spans, model response, tool request and response |
For a dropped call, start with disconnect ownership. For other failures, follow the affected component through the call path and verify its inputs before changing its configuration.
This guide combines the provider documentation cited below with diagnostic checklists. Treat timing buckets as investigation prompts and set operational thresholds from labeled calls, workload requirements, and your deployed provider configuration.
Related Guides:
- Voice Agent Incident Response Runbook - 4-Stack framework for production outages
- Debug WebRTC Voice Agents - ICE, RTP, and pipeline debugging
- How to Evaluate and Test Voice Agents - 4-Layer QA Framework
- Voice Agent Observability - 4-Layer Observability Framework
- Voice AI Latency Guide - Latency benchmarks and optimization
- Debugging Voice Agents: Logs, Missed Intents & Error Dashboards - Real-time log analysis, error dashboards, and continuous improvement loops
- OpenTelemetry for Voice Agents - OTel span hierarchies and cross-service debugging playbooks
- Failed Production Call Regression Runbook - Turn a reproduced failure into a permanent test
Why Is My Voice Bot Dropping Calls?
A voice bot drops a call when the telephony provider, SIP peer, media layer, application runtime, or caller ends the session before the intended conversation is complete. The first diagnostic job is not to guess the root cause. It is to identify the termination owner and the last healthy event.
Dropped call: A connected voice session that ends before the intended conversation is complete, excluding an intentional caller or agent hangup confirmed by call-state evidence.
Use Elapsed Time as a Fingerprint, Not a Verdict
Pair elapsed-time buckets with disconnect ownership. The windows below are triage heuristics, not universal timers: a timestamp narrows the search but does not prove the cause.
| When the call ends | First boundary to inspect | Evidence that confirms or rejects it |
|---|---|---|
| Before the agent answers | Provider routing and answer webhook | Provider status events, answer timestamp, destination, and webhook response |
| Within the first 10 seconds | Media negotiation or runtime startup | RTP and media events, codec negotiation, worker assignment, and startup errors |
| Around 20–32 seconds | Initial SIP ACK routing | SIP ladder showing whether the ACK reached the expected endpoint |
| At a repeatable longer interval | SIP session refresh | Session-Expires, refresher role, re-INVITE or UPDATE, and the final BYE |
| Immediately after sustained silence | Media loss or silence policy | Inbound RTP counters, silence event, application timeout, and disconnect reason |
| Randomly under load or after a release | Runtime lifecycle or dependency failure | Release version, worker logs, trace, resource pressure, and dependency errors |
For an inbound SIP call, Twilio error 32022 means Twilio terminated the call after waiting 32 seconds for an ACK to its 200 OK response. NAT, firewall rules, or rewritten routing headers can prevent that ACK from arriving. The 20–32 second bucket above is a triage heuristic, not a Twilio timer range; compare captures from both SIP peers to verify the signaling failure.
For longer repeatable intervals, inspect SIP session timers. RFC 4028 defines periodic session refreshes using re-INVITE or UPDATE; a failed refresh can lead a peer to terminate the session.
Identify Who Ended the Call
A transcript ending does not establish why the call ended. Identify the termination owner before attributing the failure to the model.
| Evidence | What it tells you | What it does not tell you |
|---|---|---|
| Provider “who hung up” field | Which call leg initiated termination | Why that leg chose to terminate |
| Final SIP response or BYE direction | Which SIP peer rejected or ended the dialog | Whether an upstream application triggered it |
| Runtime disconnect reason | How the agent session ended | Whether missing media originated in the network |
| Transcript ending | The last recognized speech | Who ended the call or whether audio continued |
Twilio Voice Insights call summaries document who hung up, the last SIP response, and silence detection; SIP Call-ID is available for SIP Interface and SIP Trunking calls. LiveKit SIP participant documentation distinguishes its sip.callID from the trunk provider's sip.callIDFull, exposes status in sip.callStatus, and reports the disconnect reason on the participant. Capture the equivalent fields from your provider and runtime, then join them on one timeline.
Copy This Call-Drop Evidence Packet
Create one packet per failed call. Keep raw packet captures and customer data in access-controlled systems; the diagnostic packet should contain only the identifiers and timestamps needed to retrieve them.
{ "observedAt": "2026-08-23T14:03:18.421Z", "direction": "inbound", "connectedDurationMs": 31842, "providerCallId": "provider-call-id", "sipCallId": "sip-call-id", "agentSessionId": "agent-session-id", "traceId": "trace-id", "releaseVersion": "release-sha-or-image-tag", "termination": { "initiator": "provider | sip-peer | runtime | caller | unknown", "disconnectReason": "normalized reason", "lastSipStatus": 200 }, "lastHealthyEvents": { "provider": "media flowing", "runtime": "agent speaking", "integration": "tool response received" }, "artifacts": ["provider-log", "sip-ladder", "runtime-trace"]}
Evidence packet: A joined record that lets an engineer trace one call from provider leg to SIP dialog, media session, agent runtime, dependency calls, and release version without matching by timestamp alone.
Run the Seven-Step Dropped-Call Diagnostic
- Define the failed cohort. Record direction, provider, destination, region, release version, and a precise time window.
- Plot connected duration. Look for a tight cluster around one interval before inspecting individual calls.
- Identify the termination owner. Use provider and SIP evidence before reading the transcript as a lifecycle signal.
- Join the call path. Connect provider call ID, SIP Call-ID, agent session ID, trace ID, and release version.
- Find the last healthy event. Compare provider, media, runtime, and integration timelines without forcing them into one clock.
- Reproduce one boundary at a time. Hold the scenario constant while changing only the suspected network, SIP, runtime, or dependency condition.
- Keep the fix. Convert the reproduced failure into a regression test with the same observable termination invariant.
What This Runbook Cannot Prove
A reason code is not a root cause. “Participant disconnected” or “call completed” describes lifecycle state. You still need the preceding SIP, media, runtime, and dependency events.
Packet captures have blind spots. Capture points on opposite sides of NAT, a session border controller, or a provider edge may show different routes. Compare both sides when the SIP ladder disagrees.
Not every short call is a technical drop. A caller can leave because the opening is slow, irrelevant, or confusing. When call-state evidence confirms an intentional caller hangup, move to voice agent drop-off analysis instead of debugging transport.
If You're Debugging an AI Voice Agent...
AI voice agents depend on the audio delivered by the telephony and media layers. When the symptom involves missing, delayed, or degraded audio, inspect that path alongside the ASR, LLM, and TTS traces. A transport defect can affect downstream behavior, but the symptom alone does not establish which layer failed.
| VoIP Symptom | AI Agent Impact | What Breaks |
|---|---|---|
| Elevated jitter or late packet delivery | Audio frames may miss playout deadlines | Gaps, distortion, transcription errors |
| Packet loss or loss bursts | Missing audio can impair speech recognition | Missed utterances, incomplete sentences |
| Degraded audio quality or a falling MOS estimate | Audio may become harder to understand | Review the recording and transcription together |
| NAT/firewall issues | WebRTC ICE failures, one-way audio | Agent can't hear user or vice versa |
| SIP registration failures | Calls don't connect | Complete call failure before agent loads |
| Codec mismatch | Audio format incompatibility | Garbled audio, no audio, echo |
Transport problems can propagate into the AI pipeline. Missing or delayed audio frames can surface downstream as incomplete transcripts, poor turn detection, or responses to the wrong utterance.
→ Skip to: VoIP Call Quality Checklist if you suspect network issues
Why Do Voice Agents Fail?
Voice agents fail across multiple interdependent layers: telephony, ASR, LLM orchestration, tool execution, and TTS. Single component failures cascade through subsequent decisions, making root cause diagnosis difficult. Systematic troubleshooting requires isolating whether issues stem from audio quality, semantic understanding, model failures, API integrations, or synthesis latency.
What you'll learn:
- How to identify which component (ASR, LLM, tool execution, TTS) causes specific failure patterns
- Diagnostic techniques using logs, traces, and component-level testing to isolate root causes
- Production monitoring strategies to catch issues before they impact users
Quick filter: If you're restarting services before understanding which layer failed, you're wasting time.
What Are the Common Voice Agent Failure Categories?
Voice agents combine STT (speech-to-text), NLU (natural language understanding), decision logic, response generation, and TTS. Each layer depends on previous outputs: ASR errors corrupt LLM inputs, causing downstream tool execution failures.
Failure Category Reference Table
| Category | Layer | Symptoms | Root Causes | Diagnostic Priority |
|---|---|---|---|---|
| Retrieval failures | Intelligence | Irrelevant responses, wrong facts | RAG returning wrong context | Medium |
| Instruction adherence | Intelligence | Ignoring guidelines, scope creep | Prompt drift, temperature too high | High |
| Reasoning failures | Intelligence | Logical errors, contradictions | Context overflow, model limitations | Medium |
| Tool integration | Intelligence | API errors, timeouts, wrong calls | Auth failures, parameter issues | High |
| ASR failures | Audio | Empty transcripts, wrong words | Accents, noise, phonetic ambiguity | High |
| Latency bottlenecks | Multiple | Awkward pauses, interruptions | Slow APIs, model inference, synthesis | High |
| Context loss | Intelligence | Forgetting earlier details | Token limits, state management | Medium |
| Turn-taking errors | Audio | Cutting off users, not responding | VAD misconfiguration, endpointing | High |
How Do Failures Cascade Across Layers?
Single root-cause ASR errors propagate: incorrect transcription leads to misclassified intent, which triggers wrong tool selection. External service failures cascade when slow CRM responses delay agent replies beyond the response budget established for that workflow.
| Initial Failure | Cascade Effect | User Experience |
|---|---|---|
| Network latency (Telephony) | ASR timeouts → LLM timeouts | Call drops, no response |
| ASR returning garbage (Audio) | LLM hallucinating (Intelligence) | Wrong actions, frustration |
| LLM slow (Intelligence) | Turn-taking broken | Users talk over agent |
| TTS slow (Output) | User thinks agent died | Premature hangup |
How Do You Troubleshoot ASR (Speech Recognition) Failures?
ASR Error Types and Patterns
| Error Type | Example | Root Cause | Diagnostic Check |
|---|---|---|---|
| Accent variation | "async" → "ask key" | Regional pronunciation | Test with accent datasets |
| Background noise | Random word insertions | Poor microphone, artifacts | Check audio quality scores |
| Code-mixed speech | Mixed language confusion | Multiple languages | Enable multilingual ASR |
| Low confidence | Names, numbers wrong | Critical utterance issues | Log confidence scores |
| Truncation | Sentences cut off | Aggressive endpointing | Check silence threshold |
ASR Diagnostic Checklist
- Audio reaching server? Check for audio frames in logs, verify WebRTC connection
- Codec negotiated correctly? Verify the codec, sample rate, and channel layout supported by each endpoint and the ASR input
- ASR returning transcripts? Inspect incoming audio, VAD events, and provider errors when transcripts are empty
- Confidence scores informative? If available, compare scores within the same provider/model against labeled examples; do not treat them as interchangeable accuracy measurements
- WER within your acceptance criteria? Measure against labeled transcripts for each language, acoustic condition, and critical entity type
- Provider status? Check Deepgram, AssemblyAI, Google STT status pages
ASR-Specific Fixes
- Incorporate diverse training data: accented audio, noisy environments, varied speech patterns from real production calls
- Implement noise-canceling technologies: beamforming microphones, suppression algorithms, acoustic models trained on real-world audio
- If using model-assisted transcript correction, preserve the original transcript and verify changes against the audio, especially names, numbers, and other critical details
- Deploy hardware-accelerated VAD (voice activity detection) to filter background noise before ASR processing
For detailed ASR failure patterns, see Seven Voice Agent ASR Failure Modes in Production.
If the caller hears raw provider errors, tool names, stack traces, or system warnings, use the voice agent error leakage testing checklist before treating the problem as an ordinary LLM-quality defect.
How Do You Debug LLM and Intent Recognition Failures?
LLM Failure Mode Reference
| Failure Mode | Symptoms | Root Cause | Fix |
|---|---|---|---|
| Hallucinations | Made-up facts, wrong policies | No grounding in verified data | Add RAG validation, lower temperature |
| Misclassified intent | Wrong action triggered | Ambiguous user input, poor NLU | Improve prompt, add disambiguation |
| Context overflow | Forgets earlier details | Token limit exceeded | Implement summarization, truncation |
| Cascading errors | Multiple wrong decisions | Single root mistake propagates | Add validation checkpoints |
| Rate limiting | Slow/no responses | 429 errors from provider | Implement backoff, upgrade tier |
| Prompt drift | Inconsistent behavior | Recent prompt changes | Version control prompts, A/B test |
LLM Diagnostic Checklist
- LLM endpoint responding? Direct API test, check provider status
- Rate limiting? Look for 429 errors, check tokens per minute
- Prompt changes? Review recent deployments, check for injection
- Context window? Calculate tokens per conversation, approaching limit?
- Tool calls working? Check function call logs, tool timeout errors
- Response quality? Compare to baseline, check for hallucinations
Mitigation Strategies
- Ground with verified data: integrate agents with reliable, up-to-date databases (CRM, knowledge bases, APIs)
- Implement prompt engineering: design prompts that constrain model outputs to factual, verified responses
- Evaluate model configuration changes: test supported sampling and output-length settings against a labeled task set; lower temperature alone does not guarantee factual accuracy
- Add validation checkpoints: verify critical information before executing irreversible actions
How Do You Fix Tool Execution and API Integration Failures?
Tool Call Failure Patterns
| Failure Type | Symptom | Investigation Steps | Fix |
|---|---|---|---|
| Tool not recognized | Agent continues instead of action | Check intent classification, tool definitions | Improve tool descriptions |
| Wrong tool selection | Email API called instead of SMS | Review tool descriptions, disambiguation | Add explicit tool routing |
| Parameter formatting | Tool rejects request | Validate data types, ranges, fields | Add parameter validation |
| Response misinterpretation | Incorrect follow-up actions | Check response parsing, schema validation | Fix response handling |
| Timeout | No response from tool | Check API latency, deadlines, and whether the action completed | Repair the bottleneck and verify outcome before retrying |
Tool Integration Diagnostic Steps
- Navigate to API Logs to monitor all requests/responses, check authentication errors, verify request payload structure
- Check webhook logs to verify deliveries, server response codes, timing, monitor event delivery failures
- Track tool execution results and errors through trace views showing input parameters and returned data
- Test tool integrations independently before end-to-end testing: verify API calls work outside agent context
- Measure API response latency to identify slow external services creating conversation pauses
Tool Execution Fixes
- Retry transient dependency failures with bounded backoff only when the operation is safe to repeat; verify write outcomes and enforce idempotency
- Cache frequently used data to avoid unnecessary database lookups mid-conversation
- Set external API deadlines from the conversation budget and observed dependency latency; verify the action outcome before retrying a timed-out write
- Build circuit breakers to prevent small failures from cascading into system-wide problems
How Do You Optimize TTS Latency and Quality?
TTS Performance Measurements
Compare equivalent utterances, voices, regions, and streaming modes before setting a release threshold. A full-synthesis time is meaningful only alongside output length.
| Metric | What to measure | How to use it |
|---|---|---|
| Time to first audio | Request start to the first usable audio chunk | Budget its contribution to caller-perceived response latency |
| Full synthesis | Total synthesis time and audio duration | Compare matched utterances and check whether generation keeps up with playback |
| Audio quality | Human listening labels and consistent quality measurements | Review intelligibility, artifacts, pronunciation, and prosody on representative calls |
TTS Diagnostic Methods
- Measure total latency including time-to-first-byte (TTFB) and complete audio synthesis duration
- Track component-level breakdowns to isolate delays between STT, LLM inference, and TTS generation
- Monitor tail latencies (p99) as users remember worst experiences, not average performance
- Log synthesis quality metrics: audio artifacts, volume consistency, unnatural pauses in generated speech
Optimizing TTS Performance
- Use dual streaming TTS: accepts text incrementally (token by token), begins speaking while LLM generates remaining response
- Where supported, reuse provider connections to reduce repeated setup latency
- Use the provider's supported streaming interface and verify how it handles cancellation and partial output
- Chunk long outputs at punctuation marks, stream incrementally to accelerate multi-sentence replies
For detailed latency optimization, see Voice AI Latency: What's Fast, What's Slow, and How to Fix It.
How Are Caller Hangups Different From Dropped Calls?
A caller hangup and a technical drop can produce the same final status: the call ended. Classify the event from call-state evidence before treating it as a transport incident or a conversation-quality problem.
| Observation | More consistent with caller abandonment | More consistent with a technical drop | Verify with |
|---|---|---|---|
| Caller says goodbye or declines to continue | Yes | No | Final transcript turns and caller-initiated hangup |
| Call ends at one repeatable duration across scenarios | Unlikely | Yes | Duration histogram, SIP ladder, session timer events |
| Media stops before the caller leg ends | Unlikely | Yes | RTP counters, provider media events, runtime disconnect reason |
| Caller leaves after a long pause or irrelevant answer | Yes | Possible | Turn timestamps, latency trace, caller-initiated hangup |
| Provider or SIP peer sends the final BYE | Possible | Possible | BYE direction plus the event immediately before it |
When evidence shows the caller deliberately left, investigate opening relevance, latency, interruptions, and task completion with the voice agent drop-off guide. When the provider, SIP peer, media layer, or runtime ended the session unexpectedly, keep following the technical timeline in this runbook.
VoIP Call Quality (Jitter/Packet Loss/MOS) Checklist
This section covers VoIP diagnostics that can affect AI voice agent performance. Use them when audio evidence points to the media path, and compare the same call across transport and agent traces.
Network Quality Metrics Reference
Thresholds depend on the codec, loss concealment, jitter buffer, call path, and workload. Use the provider's definitions and a known-good cohort from the same deployment; verify a suspected regression against the recording.
| Metric | Measurement | Diagnostic use |
|---|---|---|
| Packet loss | Lost RTP packets and burst patterns at a specific edge | Locate missing media and correlate it with audible gaps |
| Jitter | Variation in packet arrival timing | Inspect late packets, buffer behavior, and added playout delay |
| Latency (RTT) | Round-trip network time | Separate network delay from ASR, model, tool, and TTS delay |
| MOS estimate | Quality estimate produced by a documented method | Compare like-for-like measurements and confirm the caller impact by listening |
VoIP Diagnostic Checklist
Network & Bandwidth:
- Sufficient bandwidth? Budget for the negotiated codec, packetization, transport overhead, both directions, and peak concurrent calls
- QoS configured? Verify that voice prioritization matches your network policy across the actual media path
- Packet loss within your tested budget? Inspect RTP counters at each media edge; an ICMP ping does not measure the call's media stream
- Jitter acceptable? Check with
iperf3or RTP stream analysis - No bandwidth contention? Other applications competing during calls
NAT & Firewall:
- SIP ALG rewriting? Compare signaling on both sides of the router before changing ALG settings
- Signaling and media allowed? Verify the transport, addresses, and port ranges required by your provider and negotiated session
- ICE/STUN/TURN configured? Verify which candidate pair connects and whether the network requires a TURN relay
- Symmetric NAT handled? May require TURN relay server
- Firewall allowing RTP? Stateful inspection may block return packets
SIP & Signaling:
- SIP registration successful where required? Inspect the challenge-response sequence and final result; an initial SIP authentication challenge alone does not prove failure
- Correct SIP trunk credentials? Authentication failures = no calls
- DNS SRV records resolving? SIP often uses SRV lookups
- TLS/SRTP configured? Encryption may be required by provider
- SIP timers appropriate? Session timers, registration refresh
Codec & Audio:
- Codec negotiated correctly? Check SDP in SIP INVITE/200 OK
- Codec choice compatible? Match endpoint support and the actual media path; identify any transcoding or bandwidth tradeoffs
- Sample rate matched? Mismatch causes audio distortion
- Echo cancellation enabled? AEC required for full-duplex
- Comfort noise configured? Prevents "dead air" during silence. See the Dead Air Detection Guide for diagnosing dead-air events.
Common VoIP Issues and Fixes
| Issue | Symptoms | Diagnostic Command | Fix |
|---|---|---|---|
| SIP ALG interference | One-way audio, registration drops | Compare headers and addresses before and after the intermediary | If rewriting is the cause, test the provider-recommended ALG configuration |
| NAT traversal failure | ICE connection timeout, no audio | Check webrtc-internals ICE candidates | Configure STUN/TURN, open UDP ports |
| Codec mismatch | Garbled audio, no audio | Inspect SDP in SIP traces | Force compatible codec on both ends |
| RTP packet loss | Choppy audio, words missing | Capture the configured RTP range at both media edges | Fix the loss source; test supported FEC or buffer changes against added latency |
| DNS resolution | Intermittent call failures | Query the provider's documented SIP hostname and record type | Repair DNS or use the provider's supported routing configuration |
| TLS handshake failure | Secure calls not connecting | openssl s_client -connect sip.provider.com:5061 | Update certificates, check TLS version |
WebRTC-Specific Diagnostics
For browser-based voice agents using WebRTC:
chrome://webrtc-internals (Chrome)about:webrtc (Firefox)
Key metrics to check:
- ICE connection state: Should be "connected" or "completed"
- DTLS state: Should be "connected"
- Packets lost: Incoming/outgoing RTP packet loss
- Jitter buffer: Current delay and target delay
- Audio level: Verify audio is flowing (not 0)
RTP Stream Analysis
For deep packet inspection when standard tools don't reveal issues:
Capture RTP traffic: Replace the example range below with the ports negotiated for the call. Packet captures may contain private audio and require restricted storage.
tcpdump -i any -w voip_capture.pcap udp portrange 10000-20000
Analyze in Wireshark:
- Navigate to: Telephony → RTP → RTP Streams
- Check for packet loss percentage, jitter, delta (inter-packet timing)
- Look for sequence number gaps indicating lost packets
Key RTP metrics:
| Metric | Where to find it | What to investigate |
|---|---|---|
| Lost packets | RTP stream analysis | Loss bursts aligned with audio gaps |
| Maximum jitter | RTP stream analysis | Arrival spikes and whether the jitter buffer absorbed them |
| Mean jitter | RTP stream analysis | Changes relative to the same route and codec baseline |
| Sequence errors | RTP stream analysis | Distinguish loss, reordering, duplicates, and capture artifacts |
MOS Score Interpretation
A MOS value must be interpreted with its measurement method. A provider's network-derived estimate is not a direct measurement of ASR accuracy or task success. Twilio's Voice Insights reference describes its Voice SDK MOS estimate and provider-specific quality flags; those definitions should not be applied as universal thresholds across platforms.
Compare the same method, codec, and call path over time, then listen to representative recordings. Investigate degradation alongside packet loss, jitter, and latency, and measure transcription errors separately. A low score alone is not a reason to terminate a caller's session.
What Logging and Tracing Do You Need for Voice Agent Debugging?
Essential Logging Schema
Turn-level data (per exchange):
{ "call_id": "call_abc123", "turn_index": 3, "timestamp": "2026-01-26", "user_transcript": "I need to reschedule my appointment", "asr_confidence": 0.94, "intent": {"name": "reschedule_appointment", "confidence": 0.91}, "latency_ms": {"stt": 180, "llm": 420, "tts": 150, "total": 750}, "tool_calls": [{"name": "get_appointments", "success": true, "latency_ms": 85}], "agent_response": "I can help you reschedule..."}
Production Monitoring Essentials
Define each metric's denominator, observation window, and minimum evidence before setting an alert. Select limits from your service objectives and known-good cohorts, and keep provider/model-specific measurements separate.
| Metric | What it measures | Alert design |
|---|---|---|
| Call success rate | Calls meeting a defined technical-success condition | Compare with the service objective for that route and cohort |
| P95 end-to-end latency | Response time at the 95th percentile | Alert on a sustained breach of the workflow's response budget |
| ASR quality | Labeled transcription accuracy; provider confidence where available | Investigate drift by model, language, and acoustic condition |
| Task completion | Eligible calls achieving the intended outcome | Compare with a task-specific objective and inspect failure causes |
| Error rate | A defined class of failed calls divided by eligible calls | Alert by severity, cohort, and error budget |
Tracing Voice Agent Workflows
Tracing captures every call step: audio input, ASR output, semantic interpretation, internal prompts, model generations, tool calls, TTS output. Use OpenTelemetry for metrics, logs, traces to keep data portable across observability tools.
For detailed observability implementation, see Voice Agent Observability: The Missing Discipline.
How Do You Fix Conversation Flow and Turn-Taking Issues?
Context Loss and Memory Issues
Agents can lose important earlier conversation details when history exceeds the selected model's context budget or is truncated incorrectly. As conversations grow, critical information gets pushed out, leading to contradictions or lost problem tracking.
| Issue | Symptom | Fix |
|---|---|---|
| Token overflow | Forgets early details | Implement conversation summarization |
| State loss | Asks same question twice | Persist state externally |
| Context drift | Contradicts earlier statements | Add context anchoring prompts |
How Do You Prevent Agents from Interrupting Users?
- Tune endpointing and silence thresholds against recordings of the speakers, languages, and pause patterns your agent must support
- Use hardware-accelerated Voice Activity Detection (VAD) that handles interruptions gracefully
- Move beyond Voice Activity Detection to consider semantics, context, tone, conversational cues
- Tune VAD sensitivity based on use case: customer service needs longer thresholds than quick commands
Fixing Conversation Flow Issues
- Use hybrid context management: full server-side history for high-stakes sessions, lightweight vector summaries for general chat
- Persist critical constraints and confirm them when they change or become ambiguous; avoid requiring callers to repeat information on a fixed schedule
- Test conversation state management: verify handling of interruptions, corrections, topic changes
- Implement conversation summarization at regular intervals to maintain context within token limits
How Do You Build Error Handling and Recovery Patterns?
Resilience Design Patterns
| Pattern | Implementation | When to Use |
|---|---|---|
| Circuit breaker | Stop calling failed service | External API failures |
| Exponential backoff | Retry with increasing delays | Transient network issues |
| Graceful degradation | Fall back to simpler responses | Knowledge retrieval failures |
| Timeout limits | Deadlines based on measured dependency latency and conversation budget | Slow external services |
| Retry limits | Bounded attempts within the deadline, with idempotency and outcome checks for writes | Transient failures with a safe retry path |
User-Facing Error Recovery
- Provide clear task-level feedback: "I couldn't confirm whether that update completed. I can help check its status" instead of exposing raw errors or claiming an unverified outcome
- Build fallback logic into customer journeys: "Press 0 to speak to a live representative" when agent reaches capability limits
- Acknowledge errors transparently: "I missed that, could you repeat?" rather than guessing at misheard inputs
Continuous Improvement from Failures
- Feed production failures back into offline evaluation datasets to create continuous improvement loops
- Convert any live conversation into replayable test case with caller audio, ASR text, expected intent
- When a production call fails, preserve the relevant audio, timing, and expected outcome in a regression fixture under the applicable data controls
- Track failure resolution rates: measure time from issue identification to deployed fix
What Testing and Evaluation Strategies Work for Voice Agents?
Automated Testing Approaches
- Auto-generate test cases from agent prompts and documentation to ensure coverage
- Run representative scenarios at an authorized load that matches the capacity you need to validate, including accents, background noise, and interruptions
- Test agents in multiple languages, simulate global accents and real-world noise environments
- Implement synthetic user simulation: generate varied conversation paths to stress-test agent logic
Evaluation Metrics That Matter
| Metric category | Key metrics | Acceptance criteria |
|---|---|---|
| Conversational | Latency, interruptions, turn-taking | Budgets validated for the workflow and speaker population |
| Outcomes | Task completion, escalation rate | Task-specific goals measured on eligible scenarios |
| Quality | WER, intent accuracy, entity extraction | Labeled evaluation by language, acoustic condition, and critical field |
| Policy adherence | Data handling, required disclosures, allowed actions | Explicit requirements and release-blocking violations in the tested scenarios |
CI/CD Integration for Voice Agents
- Integrate testing into GitHub Actions, Jenkins, or CI/CD pipeline to trigger tests and block bad prompts automatically
- After each build, rerun representative scenarios and block deployment when defined task, safety, or quality criteria fail; text differing from baseline alone is not a failure
- Version control agent configurations (prompts, tools, models) alongside code for reproducible deployments
For comprehensive testing methodology, see How to Evaluate and Test Voice Agents.
Summary and Next Steps
Systematic troubleshooting requires component-level isolation: test ASR, LLM, tool execution, TTS independently before end-to-end diagnosis. Production monitoring with comprehensive logging, tracing, and observability catches issues before they impact users.
Next steps:
- Implement structured logging capturing every component's inputs, outputs, latency, confidence scores
- Set up production monitoring with alerts for latency spikes, error rate increases, quality degradation
- Build automated testing pipelines that run diverse scenarios before deployment to catch failures early
How Hamming Helps with Voice Agent Troubleshooting
Hamming helps teams keep the evidence used in this runbook connected:
- Review call audio, transcripts, evaluations, and trace context together
- Compare failures across providers, agent versions, and scenarios
- Turn a confirmed production failure into a repeatable test case
- Rerun that scenario as the agent, prompt, model, or integration changes
Provider and runtime logs remain part of the investigation. Hamming complements those systems by preserving the conversation-level evidence and regression test that show whether the correction holds.
Debug your voice agents with Hamming →
Related Guides:
- Voice Agent Drop-Off Analysis - Funnel analysis and abandonment metrics with remediation playbook
- Slack Alerts for Voice Agents - Alert templates for latency, ASR drift, jitter, and prompt regressions
- Voice Agent Monitoring KPIs - 10 production metrics with formulas and benchmarks
- Voice Agent Incident Response Runbook - Systematic debugging framework for production outages
- Voice Agent Monitoring Platform Guide - 4-Layer monitoring architecture
- Debugging Voice Agents: Logs, Missed Intents & Error Dashboards - Real-time log analysis, error dashboards, and continuous improvement loops
- Voice Agent SEV Playbook & Postmortem Template - Severity classification, response checklists, and postmortem framework for voice AI incidents

