When a user interrupts with “Not Friday, Saturday,” can the system hear the correction, stop the obsolete answer and update the date? That behavior reveals more about interaction quality than how quickly speech starts playing.
Full-duplex interaction allows input and output to occur together, with new input affecting the ongoing exchange. Streaming output alone only means playing speech progressively. A complete experience also needs ongoing listening, interruption decisions and consistent handling of the old response and task.
Full-duplex, alternating turns and streaming output
| Concept | Main concern | What it does not establish alone |
|---|---|---|
| Alternating turns | One participant responds after the other | Handling overlapping speech |
| Streaming speech output | Play generated audio incrementally | Understanding the user during playback |
| Full-duplex speech interaction | Process overlapping input and output | Correct task cancellation and recovery |
| Continuous multimodal interaction | Combine sound, scene, action and state | Availability of every modality or feature |
Full-duplex can describe observable system behavior or a model’s treatment of input and output streams. Ask which meaning a provider intends and test the corresponding behavior.
Why being interruptible is not enough
A system may pause as soon as it detects sound without distinguishing a correction from nearby conversation or speaker echo. Excessive sensitivity causes false stops; insufficient sensitivity makes it difficult to redirect the system.
After a pause, check what happens to queued audio, ongoing generation and tool calls. A tool may no longer be cancellable, or an action may already have completed. The next response should reflect the actual state.
Muting playback does not establish that the task stopped. Stopping generation does not necessarily clear audio already buffered by the output device. These behaviors need separate definitions.
Follow a correction through the task
This conceptual sequence illustrates state changes rather than any specific model architecture.
For an itinerary query, the user changes Friday to Saturday. The system recognizes a correction and updates the date. If the Friday search returns later, it should recognize the result as belonging to the obsolete task rather than append it to the new response.
Task versions, cancellation state and explicit rules for accepting results are useful design tools. The appropriate implementation depends on whether tools can be cancelled, actions can be reversed and how much waiting the application permits.
Compare speech architectures under the same conditions
A common speech pipeline combines recognition, text reasoning and synthesis. Its components can be replaced independently, but their streaming states need coordination. Other research models speech input and output directly. Kyutai’s Moshi paper describes a full-duplex approach with separate streams for user and system speech. Original Moshi paper
Architecture alone does not establish a speed or quality winner. Compare content quality, overlap behavior and end-to-end experience under the same device, network, language, context and concurrency conditions.
Measure four distinct moments
| Moment or interval | What it measures |
|---|---|
| Actual end of the user’s turn | A reference point for response waiting time |
| First audible system output | Perceived response onset |
| Interruption onset to old playback stopping | How promptly the system yields |
| Completed correction to correct task continuation | Whether new information is used |
Also record false triggers, missed interruptions, long pauses and failed recoveries. Detector latency, time to first packet and task completion time are different measurements.
Simplex SA0’s current status
Surd AI describes SA0 as a research direction connecting speech, action, expression, memory and tool use. Its public page remains a research preview. No public SA0 interaction API or release date is offered in this guide. SA0 research preview
For the narrower question of when to respond or yield, see the published MicroCF research and evaluation. A timing detector’s results do not establish the performance of the full SA0 system.
Frequently asked questions
Is full-duplex just simultaneous recognition and synthesis?
Concurrent input and output are a foundation. Practical interaction also requires handling overlap, corrections, playback and task state. Verify the behavior directly.
Does streaming TTS make a system full-duplex?
No. Streaming TTS concerns incremental output. Listening continuously and responding to new input are additional requirements.
Does full-duplex always require more GPUs?
The name alone cannot answer that. Capacity depends on the model, audio representation, concurrency, caching and scheduling. Measure usable session capacity and latency for the intended configuration.
Continue with robot interaction models and voice-agent turn-taking.