“I’d like tomorrow’s… morning flight.” If a system starts answering during that pause, a perfect transcript will not make the exchange feel natural. A voice agent needs to track whose turn it is, not just whether sound is present.
VAD provides speech-activity evidence, end-of-turn detection estimates whether a speaker has finished, and interruption detection helps decide when the system should yield. These signals are related but not interchangeable.
What do VAD, VAP, EOT and INT mean?
| Term | Main question | Important boundary |
|---|---|---|
| VAD: voice activity detection | Is speech activity present? | Silence does not necessarily end a turn |
| VAP: voice activity prediction | What activity may follow the observed conversation? | Prediction does not read future audio |
| EOT: end-of-turn detection | Is the speaker yielding the floor? | A breath may remain inside the same turn |
| INT: interruption detection | Is another participant taking the floor? | A short acknowledgment may not request a stop |
Applications can combine acoustic evidence, content and historical state. The information used by any particular model depends on its implementation and evaluation protocol, not its acronym.
Why can a fixed silence threshold fail?
A short waiting threshold may cut users off; a long one adds delay to every exchange. Speaking pace, accents, noise and habits affect how the same setting behaves.
A fixed threshold can still be a useful baseline in constrained tasks with short, predictable responses. Compare it on real pauses and corrections instead of judging speed from a few ideal recordings.
What happens after an interruption signal?
The application decides whether to pause playback, cancel old output, keep listening and modify or preserve the task. Check three levels:
- Audio: When does the old speech actually stop?
- Dialogue: Is the new utterance heard completely?
- Task: Are the old goal and pending tool results handled correctly?
If the user says “I meant Shanghai” during a weather response, pausing is only part of success. The city must also change. Counting stops alone misses whether the correction was used. State handling in full-duplex interaction
How should you evaluate turn-taking models?
Detecting more events may also increase false triggers. Report recall, false-positive rate and detection latency together. At application level, also measure premature responses, unnecessary pauses, recovery success and actual user waiting time.
Define the start and end of each timing measurement. An annotated event-to-prediction interval differs from speech onset to playback stopping. State whether computation and network time are included before comparing providers or configurations.
What does the MicroCF public record show?
The MicroCF article describes combining current activity with predictions of subsequent activity and reports separate interruption and end-of-turn results. The table below reproduces its archived TurnBench test record. It is not a claim about all deployments or today’s live ranking.
| Task | Recall | False-positive rate | Median detection latency |
|---|---|---|---|
| Interruption, INT | 98.4% | 7.3% | 747 ms |
| End of turn, EOT | 90.6% | 7.9% | 678 ms |
The public article dates its ranking check to September 23, 2026 and identifies 116 test conversations. These detector measurements are not full application response times. Development-set rescoring uses additional timing conventions and should not be mixed into this test table. MicroCF report and original-source links
Published results help explain the task and its trade-offs; they do not establish public API availability. Simplex interaction models are not publicly launched yet. Check the product pages for actual access status.
Build a more realistic voice evaluation
Include long pauses inside sentences, short pauses after complete statements, explicit corrections, acknowledgments, nearby speakers and playback echo. Label user intent, the intended handoff point and the expected application behavior. Use recordings you are authorized to evaluate.
Check both timing and task outcomes. A city correction should update the task. An “mm-hmm” may or may not call for a stop depending on context. Write down the labeling rule instead of relying on an evaluator’s changing intuition.
Frequently asked questions
Is end-of-turn detection just speech recognition?
They can be integrated, but they answer different questions. Recognition identifies what was said; turn detection concerns when to respond. A correct transcript can still arrive in a system that responds at the wrong time.
Is faster interruption always better?
Consider false triggers and recovery too. Stopping quickly on every background voice can degrade the overall experience.
Does this matter for robots?
Yes. Robots and digital humans also need to decide when to respond or listen, with additional visual-attention and action-state requirements. See the robot interaction guide.