Real-Time Intelligence Natural Conversation.
Oshara Full Duplex pairs a voice prompt with a persona prompt, set independently, in a single live session that listens and speaks at the same time.
The Tradeoff
Flexibility or timing. Never both.
Cascaded pipelines
Flexible, but waits its turn
Full-duplex systems
Fast, but fixed at build time
Oshara Full Duplex
Both, in one live session
For years, voice agents have required a tradeoff. Cascaded pipelines - speech-to-text, a language model, and then text-to-speech - give you the flexibility to choose any voice and define any role in plain language. But the agent typically has to wait until the caller finishes speaking before it can respond, making interruptions and backchannels feel unnatural or scripted.
Full-duplex systems solve the timing problem by allowing the agent to listen and speak simultaneously in the same live stream. However, they often require the voice and role to be fixed at build time.
Oshara Full Duplex is designed to eliminate that tradeoff. A voice prompt (reference_audio_url) and a persona prompt (system_prompt) can be configured independently in the same request, with both streaming into a single full-duplex session. Either can be changed in minutes without retraining the model or rebuilding the call flow, while the agent continues to listen and speak simultaneously, handle interruptions naturally, and use backchannels while the caller is still talking.
Capabilities
Capabilities
Four properties of the engine that a turn-based pipeline cannot reproduce.
Full Duplex
Caller audio and agent audio run as concurrent tracks inside the same session - not a request-then-response loop. The agent can start forming a reply before the caller finishes talking, closing the gap that makes cascaded systems feel slow.
Interruptions & Backchannel
Callers talk over the agent and it yields mid-sentence. Quiet acknowledgements land while the caller is still speaking, instead of dead air between turns.
Two Independent Prompts
reference_audio_url sets the voice. system_prompt sets the role and its boundaries. The two fields are independent - change either one in minutes, without retraining anything or rebuilding the call flow.
Built on Oshara's Voice Library
Every voice in Oshara's library - a clone or a native Nepali speaker - can front a Full Duplex persona through the same voice gallery used across the platform.
Real life usecases
The following examples show how Full Duplex behaves across different personas and languages. In every recording, Full Duplex sits in the top channel and the caller in the bottom channel.
Restaurant Reservations
EnglishRestaurant Reservations
NepaliLoan Repayment Reminder
EnglishLoan Repayment Reminder
NepaliArchitecture
Architecture
Full Duplex takes two independent inputs and keeps them independent for the life of the session:
Voice prompt
reference_audio_urla short reference clip that sets the caller-facing voice. Swap it for any voice in Oshara's library, or a fresh clone, without touching the rest of the request.
Persona prompt
system_promptplain language that sets the role, tone, and hard boundaries. Change it in minutes to turn the same voice into a different agent.
Turn-based vs. full-duplex
Most voice agents run one channel at a time - the caller waits for the agent to finish, and the agent waits for the caller. Full Duplex keeps both channels open for the whole session, which is what makes interruptions and backchannel feel natural instead of scripted.
One channel at a time. Fast, but each side still waits for a turn.
Both channels open. Speech overlaps, as it does with people.
Caller and agent audio run on independent tracks inside one live session - overlaps represent real interruptions and backchannel, not scheduling errors.
Both prompts are set once, at the start of the session, and stay in effect for its entire duration - there's no separate step to swap the voice or rewrite the persona mid-call.
How It's Built
Runs on Oshara's Platform
Full Duplex sessions use the same session and voice infrastructure as the rest of Oshara's platform - the same voice library, the same session endpoint, the same billing and usage metering. There's no separate Full Duplex stack to deploy or maintain.
Getting a Persona Live
- 1
Pick or Clone a Voice
Choose from Oshara's voice library or enroll a fresh sample - that becomes the reference_audio_url.
- 2
Write the Persona Prompt
Plain language sets the role, tone, and hard boundaries - passed straight through as system_prompt.
- 3
Go Live, Full-Duplex
The same session flow used across Oshara's platform opens one room where both sides can talk at once.
Give your agent a voice and a personality
Book a live walkthrough of Oshara Full Duplex, tailored to your use case and your voice.