Oshara Agent Builder

Build Your Own
Voice AI Agent
Deploy It Anywhere

Turn your data into a sovereign voice agent - persona, knowledge, tools, and voice, live on your site in one line of code.

Text-to-speech $0.01 / minPersona & knowledgeTools & MCP serversOne-line embedNo ML team required
Powered & Supported by
NVIDIA Inception Program memberAWS Startups member
How It Works

From Idea to Live Agent in Five Steps

1

Shape the Persona

Name it, describe it, set personality and greeting - no code.

2

Add Knowledge & Tools

Upload documents, connect URLs and tools so your agent can answer and act.

3

Pick Your Models

Choose the LLM, our Lyric or Fantom speech models, and the TTS voice that powers your agent.

4

Customize Appearance

Set the widget title, brand colors, avatar, and display settings - previewed live as you edit.

5

Deploy Anywhere

Drop one script tag on your site, or call the agent-session API. The widget appears automatically.

Everything You Can Configure

A Full Studio for Agentic Voice

Persona, knowledge, tools, models, appearance, and forms - tuned in one place, with a live preview of your widget the whole way.

Persona Designer

Tone, personality traits, and greeting style - every conversation sounds like your brand.

Knowledge Base

PDFs and URLs, chunked and retrieved automatically mid-conversation.

Tools & MCP Servers

MCP servers, HTTP endpoints, and client events - your agent can act, not just talk.

Model Control

Any LLM, our Lyric (90+ languages) or Fantom (1600+ languages) speech models, and a TTS voice library.

Branded Appearance

Name, colors, avatar, and watermark - previewed live as you edit.

Multi-Step Forms

Forms your agent surfaces mid-call to book demos, capture leads, or run intake.

Voice Library

Pick a Voice That Fits Your Brand

Production voices tuned for English, Nepali, and Hindi - preview each one before you assign it to an agent.

Sign up
JamesComposed

Calm and composed tone with quiet authority.

Under the Hood

Built for Real-World Constraints

Sovereign, low-latency, and fine-tuned on your data - the infrastructure behind every agent.

Sovereign Security

All audio and data processed locally - nothing sent to foreign clouds.

Low Latency

Optimized for real-time voice UX, production-ready out of the box.

Language & Dialect Precision

Fine-tuned on Nepali and regional dialects - nuance, grammar, and vocabulary included.

Sentiment-Adaptive Tone

The voice model adapts tone to emotional context, reducing churn.

Fine-Tuning Automation

OFTA handles data prep, training, and evaluation - bring your own data, we run the pipeline.

Information & Call Transfer

Hand Off Conversations Between Agents

Wire up call-transfer rules so a conversation moves - with full context - to the right specialist.

Call Transfer Settings

Decide when a conversation escalates, then preview the hand-off.

Allow agent-to-agent transfer
Carry conversation context
Auto-escalate on intent match
Support Bot
First responder
Billing Agent
Payments & invoicing
Tech Support
Troubleshooting
Live Human
Senior escalation

One agent routes a live conversation out to the right specialist.

Full Duplex

Real Conversation, Both Ways at Once

Your agent listens and speaks in the same live session - voice and persona set independently, so it never waits its turn.

Turn-based todayTypical agents
Caller
Full Duplex

One channel at a time. Fast, but each side still waits for a turn.

Full-duplexOshara
Caller
Full Duplex

Both channels open. Speech overlaps, as it does with people.

AgentUser

Restaurant Reservations

English
0:00.00 / 0:24.02

Loan Repayment Reminder

Nepali
0:00.00 / 0:29.06
Actions

Give Your Agent a Form to Work With

Pick a template or build your own — your agent surfaces the right form mid-call, fields pre-filled from the conversation.

General Inquiry
Get in Touch
I'll make sure the right person sees this right away.
Name
Sam Rivera
Email
sam@email.com
Phone number
+1 (555) 204-8821
Team size
11–50
Message
Describe your question…
Choose a template

Ready-made forms, live in the widget.

Click a template to preview it - your agent picks automatically mid-call.

Forms that fill themselves - detected mid-call, pre-filled, ready to confirm.

Deploy on Your Site

Your Agent. Your Website. Two Ways to Ship.

Embed the hosted widget with one snippet, or build a fully custom experience on top of the agent-session API.

No code

Hosted Widget

We host it - paste one script tag on any page and the widget appears.

<script
  src="https://api.oshara.ai/widget.js"
  data-agent="your-agent-id">
</script>
Code

Agent-Session API

Mint a session with your API key and connect your own UI to the live agent.

curl -X POST https://api.oshara.ai/api/agents/agent-session/ \
  -H "x-api-key: sk_..." \
  -H "Content-Type: application/json" \
  -d '{ "agent": "your-agent-id" }'
Lock it to your domains. Switch languages on the fly.

Every conversation stays on infrastructure you control.

Read the docs
Pricing

One Cent a Minute of Speech

You pay for the audio your agent actually generates - a flat $0.01 per minute of text-to-speech. No seats, no platform fee, no idle-time billing.

Text-to-speech
$0.01/ min

That's $0.60 an hour of generated speech - the same rate whether you ship one agent or a hundred.

  • No seats, no platform fee
  • Same flat rate on every voice
Estimate your bill100 min
$1.00total

100 min × $0.01 per minute

1.7 hours of generated audio

100 min100k min
Start building free

Final usage is billed per minute generated.

Ready to Build Your Agent?

Design the persona, feed it your knowledge, and deploy to your site in an afternoon - no ML team required.