← Models & interaction guides
GUIDES / Voice interaction basics

When should a voice agent respond? VAD, turn-taking and interruptions

Learn the differences between voice activity, end-of-turn and interruption detection, how to measure timing, and what Simplex MicroCF’s public results show.

Surd AI · Research & engineeringEnglish

“I’d like tomorrow’s… morning flight.” If a system starts answering during that pause, a perfect transcript will not make the exchange feel natural. A voice agent needs to track whose turn it is, not just whether sound is present.

VAD provides speech-activity evidence, end-of-turn detection estimates whether a speaker has finished, and interruption detection helps decide when the system should yield. These signals are related but not interchangeable.

What do VAD, VAP, EOT and INT mean?

Term Main question Important boundary
VAD: voice activity detection Is speech activity present? Silence does not necessarily end a turn
VAP: voice activity prediction What activity may follow the observed conversation? Prediction does not read future audio
EOT: end-of-turn detection Is the speaker yielding the floor? A breath may remain inside the same turn
INT: interruption detection Is another participant taking the floor? A short acknowledgment may not request a stop

Applications can combine acoustic evidence, content and historical state. The information used by any particular model depends on its implementation and evaluation protocol, not its acronym.

Why can a fixed silence threshold fail?

A short waiting threshold may cut users off; a long one adds delay to every exchange. Speaking pace, accents, noise and habits affect how the same setting behaves.

A fixed threshold can still be a useful baseline in constrained tasks with short, predictable responses. Compare it on real pauses and corrections instead of judging speed from a few ideal recordings.

What happens after an interruption signal?

The application decides whether to pause playback, cancel old output, keep listening and modify or preserve the task. Check three levels:

  1. Audio: When does the old speech actually stop?
  2. Dialogue: Is the new utterance heard completely?
  3. Task: Are the old goal and pending tool results handled correctly?

If the user says “I meant Shanghai” during a weather response, pausing is only part of success. The city must also change. Counting stops alone misses whether the correction was used. State handling in full-duplex interaction

How should you evaluate turn-taking models?

Detecting more events may also increase false triggers. Report recall, false-positive rate and detection latency together. At application level, also measure premature responses, unnecessary pauses, recovery success and actual user waiting time.

Define the start and end of each timing measurement. An annotated event-to-prediction interval differs from speech onset to playback stopping. State whether computation and network time are included before comparing providers or configurations.

What does the MicroCF public record show?

The MicroCF article describes combining current activity with predictions of subsequent activity and reports separate interruption and end-of-turn results. The table below reproduces its archived TurnBench test record. It is not a claim about all deployments or today’s live ranking.

Task Recall False-positive rate Median detection latency
Interruption, INT 98.4% 7.3% 747 ms
End of turn, EOT 90.6% 7.9% 678 ms

The public article dates its ranking check to September 23, 2026 and identifies 116 test conversations. These detector measurements are not full application response times. Development-set rescoring uses additional timing conventions and should not be mixed into this test table. MicroCF report and original-source links

Published results help explain the task and its trade-offs; they do not establish public API availability. Simplex interaction models are not publicly launched yet. Check the product pages for actual access status.

Build a more realistic voice evaluation

Include long pauses inside sentences, short pauses after complete statements, explicit corrections, acknowledgments, nearby speakers and playback echo. Label user intent, the intended handoff point and the expected application behavior. Use recordings you are authorized to evaluate.

Check both timing and task outcomes. A city correction should update the task. An “mm-hmm” may or may not call for a stop depending on context. Write down the labeling rule instead of relying on an evaluator’s changing intuition.

Frequently asked questions

Is end-of-turn detection just speech recognition?

They can be integrated, but they answer different questions. Recognition identifies what was said; turn detection concerns when to respond. A correct transcript can still arrive in a system that responds at the wrong time.

Is faster interruption always better?

Consider false triggers and recovery too. Stopping quickly on every background voice can degrade the overall experience.

Does this matter for robots?

Yes. Robots and digital humans also need to decide when to respond or listen, with additional visual-attention and action-state requirements. See the robot interaction guide.

SIMPLEX CD / API

Try a decision with your own input.

Create an account, explore decision models in the Console, or get started with the API documentation.