Gnani, Sarvam Bulbul V3, and Murf Falcon Join the Vomyra Voice Model Lineup | Vomyra Docs
Updates
New FeatureVoice
Gnani, Sarvam Bulbul V3, and Murf Falcon Join the Vomyra Voice Model Lineup
Vomyra now includes Gnani's enterprise-grade Indian STT, Sarvam Bulbul V3's 35+ Indic voices across 11 languages, and Murf Falcon's sub-55ms ultra-low-latency TTS making Vomyra the only AI voice agent platform in India that runs every major frontier speech model under one roof.
What's new
Vomyra has added three new speech providers to the platform: Gnani (automatic speech recognition for Indian enterprises), Sarvam Bulbul V3 (Indic-first text-to-speech across 11 Indian languages), and Murf Falcon (the world's lowest-latency streaming TTS API at sub-55ms model latency).
This brings Vomyra's live, production-ready voice model roster to the most complete lineup available on any single Indian AI voice agent platform β combining frontier Indian-language speech models with global leaders in a single, switchable stack.
Why this matters for your agents
Every business in India is different. A real estate agent qualifying 99acres leads in Noida needs different voice characteristics than a BFSI lender calling borrowers across rural Maharashtra, or an EdTech brand reaching NEET aspirants in Tamil Nadu. Until now, accessing the right speech model for each use case meant maintaining separate API contracts, separate infrastructure, and separate agent deployments across multiple vendors.
On Vomyra, you switch voice models per agent, per campaign, or per use case β in one platform, without rebuilding your workflow.
Gnani β Enterprise-Grade Indian Speech Recognition
What Gnani is ?
Gnani.ai is a Bengaluru-built voice AI company founded in 2016, specializing in automatic speech recognition (ASR) for Indian enterprises. Its ASR engine is trained on 14 million hours of real telephonic audio across 20+ Indian languages and dialects, including seamless mid-sentence code-switching between Hindi and English (Hinglish), Tamil-English (Tanglish), and other mixed-language patterns common in Indian business calls.
Gnani processes over 30 million voice interactions daily and is trusted by India's largest BFSI, telecom, and healthcare enterprises β including HDFC Bank, Airtel, and Tata Motors. Its ASR handles regional accents, low-bandwidth call audio, background noise, and rapid code-switching that generic global ASR models consistently struggle with.
What Gnani unlocks on Vomyra
Pair Gnani ASR with any TTS voice in Vomyra's stack to build agents that understand Indian speech accurately β even in noisy call center environments, even when callers switch mid-sentence between Hindi and English.
Best for:
BFSI outreach: loan qualification calls, EMI reminder calls, insurance lead follow-up
Healthcare: patient intake, appointment confirmation, prescription reminders in regional languages
Government and regulated industries: applications requiring on-premise-grade ASR accuracy
Any use case where your callers speak Hinglish, Tanglish, or regional-accented Hindi
Languages supported: Hindi, Bengali, Tamil, Telugu, Kannada, Malayalam, Marathi, Gujarati, Punjabi, Odia, and 10+ more Indian languages with code-switching support.
Sarvam Bulbul V3 β India's Sovereign TTS, Now on Vomyra
What Sarvam Bulbul V3 is
Sarvam AI is a Bengaluru AI research company building India's sovereign AI stack, founded by Vivek Raghavan and Pratyush Kumar from the AI4Bharat initiative at IIT Madras. In April 2025, Sarvam was selected under India's IndiaAI Mission to build foundational models with government-backed GPU capacity. In June 2026, Sarvam became a unicorn β a $1.5B valuation Series B led by HCLTech.
Bulbul V3 is Sarvam's third-generation text-to-speech model, purpose-built for Indian languages. In a Josh Talks blind listening study with 20,000+ votes, Bulbul V3 was rated the most natural-sounding TTS for Indian languages β outranking global models including Google, Amazon, and ElevenLabs on Indian language naturalness.
Bulbul V3 β technical specifications
Capability Bulbul V3
Languages 11 Indian languages
Voices 35+ professional voice artists
Audio quality 48kHz full-band output
Latency Sub-250ms first-byte via WebSocket streaming
Code-switching Native Hinglish / Tanglish support
Architecture LLM-based prosody analysis before audio generation
The prosody difference: Hindi stress patterns, pitch contours, and sentence intonation are fundamentally different from English. Global TTS models trained on English apply English prosody to Hindi text β the result is intelligible but sounds wrong to native speakers. Bulbul V3 uses an LLM-based text analysis layer to infer emphasis, pauses, and pacing before generating audio, producing speech that sounds like it was spoken by a native Indian speaker, not translated by a machine.
What Sarvam Bulbul V3 unlocks on Vomyra
Agents using Bulbul V3 speak with the cadence, intonation, and warmth of a native Indian speaker β delivering measurably higher connection rates, longer call engagement, and better conversion on outreach campaigns targeting Hindi-first, regional-language, and code-mixed audiences.
Best for:
Real estate outreach in Hindi, Marathi, Gujarati, and Tamil
EdTech counseling calls to NEET/JEE/MBA aspirants in regional languages
D2C and FMCG brand campaigns targeting Tier 2 and Tier 3 cities
Any outbound campaign where the buyer is more comfortable in their mother tongue than English
Murf Falcon β The World's Fastest Streaming TTS, Now on Vomyra
What Murf Falcon is?
Murf Falcon is a streaming text-to-speech model from Murf AI, engineered specifically for real-time conversational AI agents. In third-party benchmarks across 33 global locations using apiping.io, Falcon outperformed ElevenLabs, OpenAI, Cartesia, and Deepgram on time-to-first-audio β confirming its position as the lowest-latency production TTS API available in 2026.
Murf Falcon β technical specifications
Capability Murf Falcon
Model latency Sub-55ms
Time-to-first-audio ~100ms end-to-end
Languages 35 languages, 150+ voices
Concurrency 10,000+ concurrent calls with no latency degradation
Pricing $0.01 per minute
Architecture Compute-efficient NLP with disentangled phoneme encoding
Compliance SOC 2 Type II Β· ISO 27001 Β· GDPR Β· HIPAA
Why latency matters in voice AI: In a real-time conversation, every millisecond of TTS latency is perceived as hesitation β the agent sounds like it's "thinking." At sub-55ms model latency, Murf Falcon produces responses that feel immediate and natural. At 130ms+ (where most other TTS models operate), callers notice the delay. For outbound sales calling, faster response = more natural conversation = higher conversion.
The concurrency advantage: Murf Falcon maintains sub-55ms latency even at 10,000 simultaneous calls. For Vomyra customers running large outbound campaigns β real estate developers calling 5,000 99acres leads in a morning, or FMCG brands running pan-India distributor outreach β Falcon ensures every caller gets the same instant-response experience, regardless of campaign volume.
What Murf Falcon unlocks on Vomyra
Pair Falcon with Gnani ASR or Sarvam Bulbul V3 for a fully India-optimized speech stack β or use Falcon for English, international, or global campaigns where raw speed and voice quality are the primary requirements.
Best for:
High-volume outbound campaigns (5,000+ simultaneous calls)
International or NRI outreach in English
SaaS and technology companies where voice quality and agent responsiveness reflect product brand
Any scenario where real-time, human-paced conversation speed is non-negotiable
Vomyra Agent Assist β Your AI Co-Pilot on Every Live Call
What Agent Assist is
While Vomyra's six AI Sales Agents (Research, Outreach, Qualification, Closing, Follow-Up, Quotation) handle fully autonomous calls without a human on the line, Vomyra Agent Assist is designed for the moments when a human sales rep is on the call and needs an AI co-pilot in real time.
Agent Assist listens to every live call as it happens and surfaces β silently, to the rep only β the right information at the right moment. Script prompts. Objection rebuttals. Product details. CRM data from previous interactions. Pricing information. Competitive comparisons. Compliance flags. Next best action recommendations. All of it, live, before the rep even knows they need it.
What Agent Assist does during a live call
Real-time battle cards: When a prospect mentions a competitor β "We're also looking at Bolna" or "What's different from what Gnani offers?" β Agent Assist instantly surfaces a battle card with specific, factual differentiators tailored to that competitor. No more fumbling. No more generic responses.
Live objection handling: When a prospect says "Send me details first" or "My manager needs to decide," Agent Assist suggests the most effective response based on your team's historical conversion data β not a generic script, but the rebuttal that actually works for your specific product and buyer profile.
CRM surfacing: If the prospect was in your CRM six months ago, Agent Assist surfaces the full history β what they said, what they were offered, what objection they raised, what follow-up was promised. The rep knows everything without pausing to look it up.
Compliance guardrails: For BFSI, healthcare, and regulated industries, Agent Assist flags in real time when a conversation is approaching a compliance boundary β before the rep says something they shouldn't.
Live coaching: Post-call, Agent Assist generates a coaching summary for every rep β what they did well, where they deviated from best practice, and what the top performers say differently in the same scenario.
Why Agent Assist is different from a script
A script tells a rep what to say. Agent Assist tells a rep what to say right now, based on what the customer just said, what the rep's CRM history shows, and what has worked in the past. It is reactive intelligence, not static guidance.
The difference in outcome: reps using Agent Assist handle objections 40% faster, spend 60% less time searching for product information during calls, and close at measurably higher rates because they are never caught unprepared.
Best for
Enterprise sales teams where deal size justifies human involvement but AI can meaningfully improve conversion
BFSI contact centers handling loan applications, insurance queries, or wealth management conversations
Real estate developers with senior sales consultants handling site visit follow-up calls
EdTech counselors managing high-stakes enrollment conversations with parents
Any inbound or outbound team where human reps are on the line but call consistency and conversion rate need to improve
Vomyra is now the only platform with all of this in one place
This update completes what no other voice AI platform in India β or globally β offers in a single subscription:
Voice Model Category What it's best for
AWS Nova 2 Sonic Realtime voice Ultra-low latency English/multilingual
OpenAI GPT Realtime (gpt-realtime-mini) Realtime LLM voice Advanced reasoning + voice in one model
Azure Voice Live Realtime voice Enterprise Microsoft stack integration
Cartesia Voice cloning Clone your brand voice or sales rep voice
ElevenLabs Multilingual TTS Premium voice quality, global languages
xAI Grok Voice Realtime voice Latest frontier voice model
Gnani ASR (new) Indian STT 14M hours Indian telephonic training, 20+ languages
Sarvam Bulbul V3 (new) Indian TTS 35+ voices, 11 Indian languages, Hinglish native
Murf Falcon (new) Streaming TTS Sub-55ms latency, 10K concurrent calls
Every model is available on every Vomyra plan. Switch per agent, per campaign, or per use case β without rebuilding your workflow, without a new API contract, and without a new vendor relationship.
How to enable Gnani, Sarvam Bulbul V3, or Murf Falcon on your agents
Log in to your Vomyra dashboard at console.vomyra.com
Open any existing agent or create a new one
Navigate to Voice Settings β Speech Provider
Select Gnani (for STT), Sarvam Bulbul V3 (for TTS), or Murf Falcon (for TTS)
Choose your preferred voice (for Sarvam and Murf) or language configuration (for Gnani)
Save and test with a live call
No additional API keys required. No separate billing. All three models are included in your existing Vomyra plan.