Persona Engine

Real-Time Intelligence Natural Conversation.

Oshara Full Duplex pairs a voice prompt with a persona prompt, set independently, in a single live session that listens and speaks at the same time.

01

The Tradeoff

Flexibility or timing. Never both.

Cascaded pipelines

Flexible, but waits its turn

Full-duplex systems

Fast, but fixed at build time

Oshara Full Duplex

Both, in one live session

For years, voice agents have required a tradeoff. Cascaded pipelines - speech-to-text, a language model, and then text-to-speech - give you the flexibility to choose any voice and define any role in plain language. But the agent typically has to wait until the caller finishes speaking before it can respond, making interruptions and backchannels feel unnatural or scripted.

Full-duplex systems solve the timing problem by allowing the agent to listen and speak simultaneously in the same live stream. However, they often require the voice and role to be fixed at build time.

Oshara Full Duplex is designed to eliminate that tradeoff. A voice prompt (reference_audio_url) and a persona prompt (system_prompt) can be configured independently in the same request, with both streaming into a single full-duplex session. Either can be changed in minutes without retraining the model or rebuilding the call flow, while the agent continues to listen and speak simultaneously, handle interruptions naturally, and use backchannels while the caller is still talking.

02

Capabilities

Capabilities

Four properties of the engine that a turn-based pipeline cannot reproduce.

01

Full Duplex

Caller audio and agent audio run as concurrent tracks inside the same session - not a request-then-response loop. The agent can start forming a reply before the caller finishes talking, closing the gap that makes cascaded systems feel slow.

02

Interruptions & Backchannel

Callers talk over the agent and it yields mid-sentence. Quiet acknowledgements land while the caller is still speaking, instead of dead air between turns.

03

Two Independent Prompts

reference_audio_url sets the voice. system_prompt sets the role and its boundaries. The two fields are independent - change either one in minutes, without retraining anything or rebuilding the call flow.

04

Built on Oshara's Voice Library

Every voice in Oshara's library - a clone or a native Nepali speaker - can front a Full Duplex persona through the same voice gallery used across the platform.

Real life usecases

The following examples show how Full Duplex behaves across different personas and languages. In every recording, Full Duplex sits in the top channel and the caller in the bottom channel.

AgentUser

Restaurant Reservations

English
0:00.00 / 0:24.02

Restaurant Reservations

Nepali
0:00.00 / 0:34.34

Loan Repayment Reminder

English
0:00.00 / 0:31.34

Loan Repayment Reminder

Nepali
0:00.00 / 0:29.06
04

Architecture

Architecture

Full Duplex takes two independent inputs and keeps them independent for the life of the session:

Voice prompt

reference_audio_url

a short reference clip that sets the caller-facing voice. Swap it for any voice in Oshara's library, or a fresh clone, without touching the rest of the request.

Persona prompt

system_prompt

plain language that sets the role, tone, and hard boundaries. Change it in minutes to turn the same voice into a different agent.

Turn-based vs. full-duplex

Most voice agents run one channel at a time - the caller waits for the agent to finish, and the agent waits for the caller. Full Duplex keeps both channels open for the whole session, which is what makes interruptions and backchannel feel natural instead of scripted.

Turn-based todayTypical agents
Caller
Full Duplex

One channel at a time. Fast, but each side still waits for a turn.

Full-duplexOshara
Caller
Full Duplex

Both channels open. Speech overlaps, as it does with people.

Caller Audio
Agent Audio
Voice Prompt
Persona Prompt
Live Session

Caller and agent audio run on independent tracks inside one live session - overlaps represent real interruptions and backchannel, not scheduling errors.

Both prompts are set once, at the start of the session, and stay in effect for its entire duration - there's no separate step to swap the voice or rewrite the persona mid-call.

How It's Built

Runs on Oshara's Platform

Full Duplex sessions use the same session and voice infrastructure as the rest of Oshara's platform - the same voice library, the same session endpoint, the same billing and usage metering. There's no separate Full Duplex stack to deploy or maintain.

Getting a Persona Live

  1. 1
    Pick or Clone a Voice

    Choose from Oshara's voice library or enroll a fresh sample - that becomes the reference_audio_url.

  2. 2
    Write the Persona Prompt

    Plain language sets the role, tone, and hard boundaries - passed straight through as system_prompt.

  3. 3
    Go Live, Full-Duplex

    The same session flow used across Oshara's platform opens one room where both sides can talk at once.

Give your agent a voice and a personality

Book a live walkthrough of Oshara Full Duplex, tailored to your use case and your voice.