Next-Gen Voice AI

Every languagedeserves a voice.

Most voice AI works well in a handful of languages. We started with one the industry had ignored, and we're building toward all of them.

24%
Word error rate, Nepali
from 48%
200 ms
To first audio
from 1,100 ms
Lower cost per minute
vs. hosted voice APIs

Oshara figures are our own measurements. Baselines are indicative of typical off-the-shelf performance on Nepali a reference point, not a benchmark of any named product.

01Why it matters

Voice AI is exploding. For a handful of languages.

Voice AI is one of the fastest-growing categories in software. Almost all of that growth is being built for a handful of languages, exactly the gap we started in.

$2.4B → $47.5B

Voice AI agents market, 2024 → 2034 (34.8% CAGR)

$11.6B → $41.4B

Conversational AI market, 2024 → 2030

32M

Nepali speakers worldwide, a market with no high-quality voice AI until now

Sources: Market.us, Voice AI Agents Market (2024); Grand View Research, Conversational AI Market Report. Third-party market projections, not Oshara measurements.

02Our direction

We started with Nepali. The mission is every language.

Nepali has roughly 32 million speakers and, until recently, no high-quality voice AI. We built our own speech models to fix that, and we're using the same method to build toward English, Spanish and beyond. The numbers below are the proof; the audio underneath is the same claim, out loud.

What changed
Off-the-shelfFine-tuned
Generic recognitionTrained on our languages
Robotic deliveryNatural prosody
Mispronounced namesCorrect pronunciation
Slow first responseFaster inference
Research qualityProduction quality
What it measures
Off-the-shelf baselineOshara fine-tuned
Word error rate (Nepali)50%
48%24%
Time to first audio82%
1,100ms200ms
Mean opinion score+42%
2.9 / 54.12 / 5

Oshara figures are our own measurements. Baseline figures are indicative of typical off-the-shelf performance on Nepali a reference point, not a benchmark of any named product.

Hear it for yourself

The same passage, the same voice, before and after training on Nepali.

Nepali sample

आजका प्रमुख समाचारअनुसार देशभर मनसुनी वर्षा सक्रिय रहँदा केही क्षेत्रमा दैनिक जनजीवन प्रभावित भएको छ। यसैबीच, कृषि तथा पूर्वाधार विकासका लागि करिब रु. २५ करोड बराबरका नयाँ कार्यक्रम सञ्चालनमा ल्याइएको जनाइएको छ। भारततर्फ नेपालको विद्युत् निर्यात ६० मेगावाट पुगेको छ भने पर्यटन क्षेत्रमा यस वर्ष १२ लाखभन्दा बढी विदेशी पर्यटक भित्र्याउने लक्ष्यअनुसार आगमन क्रमशः बढिरहेको बताइएको छ। साथै, डिजिटल सेवा तथा अनलाइन कारोबारको विस्तारसँगै विभिन्न सरकारी र निजी सेवाहरू थप सहज रूपमा उपलब्ध हुन थालेका छन्।

Before fine-tuning
After fine-tuning
03Try it yourselfLive

It speaks, and it listens

Our models, running live - pick a language, type, generate, or speak, and experience it in action.

251/600
04One platform, multiple languages

One method, a family of voices

Same training method, same platform, different language. Nepali is where we've invested the most

Voice

Ladies and gentlemen… [clear throat] the moment we've all been waiting for has finally arrived!

After months of relentless preparation, [gasp] late nights, and endless dedication, today is the big day.

I'll admit… sleep wasn't exactly easy last night. [chuckle] The excitement was just too much to handle.

[cough]Along the way, we faced challenges that tested our patience and determination.

Some obstacles honestly felt impossible at the time… [groan] like there was no way through.

But [sniff] we refused to quit.

[shush] And now… here we are, standing on the edge of something truly incredible.

05How it works

Understands naturally. Thinks acts and Responds like a person.

Speech in, reasoning and tools in the middle, speech back out, fast enough to feel like a real conversation. Everything below is the real product, running live.

Search its knowledge base

Answers grounded in uploaded documents and crawled websites.

Search the web

Live lookup when the answer isn't in its own sources.

Send email

Dispatch a message as part of the conversation.

Call your API

Custom HTTP tools with typed inputs, hitting your endpoints.

Connect over MCP

Attach an MCP server and inherit its whole toolset.

Capture structured data

Fill and submit a multi-step form mid-call, then POST it.

Hand off to a specialist

Transfer the conversation to another agent, or escalate to a live human - with full context carried over.

Fire a browser event

Trigger anything in the host page from inside the call.

06Human handoff

When it's time for a real person

The agent knows when a request is outside its scope - and hands off everything it knows, so the caller never repeats themselves.

07See it end to endLive

Build it, call it, ship it

The same agent, two ways to experience it - call it live, or hand it a form to fill by voice. When it's ready, deploy it with one script tag.

Click to switch demos

index.htmlDeploy in minutes
<script src="https://api.oshara.ai/widget.js" data-agent="your-agent"></script>
What's next

This is not another chatbot. It is the future of voice-first AI.

  1. We started

    by improving speech recognition and voice synthesis for a language nobody had built for.

  2. Today

    we run an end-to-end platform: our own STT and TTS, an LLM that acts, dynamic forms, and no-code voice agents you deploy with one script tag.

  3. Next

    the same method applied to English, Spanish and beyond, and duplex conversation, where people and agents speak at the same time, the way people actually do.

Now go make one that sounds like your customers

Describe how it should sound, give it what it needs to know, pick a voice, embed it.