Full-duplex audio
Accepts streaming audio input and returns incremental spoken output.
AI Models
Alibaba Cloud's end-to-end Qwen Omni realtime speech model for live phone agents. Stream call audio in and play spoken replies back in the same session.
Qwen Omni Realtime is Alibaba Cloud's end-to-end speech model for live conversations. It supports continuous spoken dialogue, turn detection modes, and function tools.
With AgentDuet, the phone or WhatsApp connection stays on AgentDuet while your application bridges call audio to Model Studio and plays spoken replies back.
For setup, credentials, and API details, use the Documentation link above.
The technical characteristics that matter when deciding whether this provider belongs in your application.
Accepts streaming audio input and returns incremental spoken output.
Supports server voice activity detection, semantic turn detection, and push-to-talk style control.
Requests approved application functions with structured arguments.
Selects system or cloned voices for generated speech responses.
Practical scenarios where the provider's role is clear and the surrounding systems remain under application control.
Run continuous spoken conversations with low-latency replies.
Answer routine requests and resolve complex inquiries with collected context.
Serve callers with Qwen Omni realtime speech models suited to language coverage needs.
These providers occupy a similar role in the stack. Compare their strengths, operating model, and surrounding services.
AgentDuet handles the phone and messaging boundary. Your model runs the conversation.