Qwen Omni logo

AI Models

Qwen Omni

Alibaba Cloud's end-to-end Qwen Omni realtime speech model for live phone agents. Stream call audio in and play spoken replies back in the same session.

Speech-to-speechRealtimeAlibaba CloudMultilingualFunction tools
Documentation
Where it fits

Qwen Omni in your voice stack.

  1. TelephonyPhone networkInbound call
  2. AgentDuetCall + audioAnswer, stream
  3. Speech-to-speechQwen OmniListens, call functions, replies
  4. ToolsActionsCRM, calendar, APIs
Overview

What it contributes to the stack.

Qwen Omni Realtime is Alibaba Cloud's end-to-end speech model for live conversations. It supports continuous spoken dialogue, turn detection modes, and function tools.

With AgentDuet, the phone or WhatsApp connection stays on AgentDuet while your application bridges call audio to Model Studio and plays spoken replies back.

For setup, credentials, and API details, use the Documentation link above.

Capabilities

The technical characteristics that matter when deciding whether this provider belongs in your application.

Full-duplex audio

Accepts streaming audio input and returns incremental spoken output.

Turn detection modes

Supports server voice activity detection, semantic turn detection, and push-to-talk style control.

Function tools

Requests approved application functions with structured arguments.

Configurable voices

Selects system or cloned voices for generated speech responses.

Common use cases

Practical scenarios where the provider's role is clear and the surrounding systems remain under application control.

Voice assistants

Run continuous spoken conversations with low-latency replies.

Contact center automation

Answer routine requests and resolve complex inquiries with collected context.

Multilingual support

Serve callers with Qwen Omni realtime speech models suited to language coverage needs.

Build a voice agent.

AgentDuet handles the phone and messaging boundary. Your model runs the conversation.

Explore the SDK