Bidirectional audio
Receives caller audio and emits generated speech concurrently over a live session.
AI Models
Google's real-time speech model for live phone agents. Stream call audio in and play spoken replies back in the same session.
Gemini Live is Google's real-time speech model built for live conversations. It listens and speaks in the same session, supports interruptions, and can request tools without dropping the call. That makes it a strong fit when you want a phone agent that sounds continuous, not turn-based IVR.
On AgentDuet, your application bridges live call audio to a Gemini Live session and plays speech back to the caller while AgentDuet handles answer and hangup.
For setup, credentials, and API details, use the Documentation link above.
The technical characteristics that matter when deciding whether this provider belongs in your application.
Receives caller audio and emits generated speech concurrently over a live session.
Reports interrupted responses so queued AgentDuet playback can be cleared.
Requests approved application functions without leaving the conversation.
Can combine audio with text and supported visual input supplied by the application.
Practical scenarios where the provider's role is clear and the surrounding systems remain under application control.
Collect requirements, check availability through tools, and confirm a next step by voice.
Handle routine questions while routing account actions through controlled systems.
Conduct an adaptive conversation and capture structured outcomes for follow-up.
These providers occupy a similar role in the stack. Compare their strengths, operating model, and surrounding services.
AgentDuet handles the phone and messaging boundary. Your model runs the conversation.