All stories
AI

Sarvam AI Unveils Saaras V4: A Groundbreaking Multilingual Speech-to-Text Model for All 22 Indian Languages and Global English

Sarvam AI's Saaras V4 model accurately transcribes and translates all 22 official Indian languages alongside global English, dramatically lowering digital barriers and fostering inclusion across India's diverse linguistic landscape.

By TECH NEWS Editorial·Source:MarkTechPost·3 min read·2h ago

✨ This content was summarized and interpreted by AI; it may contain errors — please verify accuracy with the original sources. Learn more

Share

Listen to this story

0:00 / 0:00
Sarvam AI Unveils Saaras V4: A Groundbreaking Multilingual Speech-to-Text Model for All 22 Indian Languages and Global English

Sarvam AI has introduced Saaras V4, a groundbreaking speech-to-text model designed to accurately transcribe and translate all 22 official Indian languages alongside global English, marking a significant leap in multilingual AI accessibility. This Bengaluru-based company's latest offering integrates an audio encoder with a 3-billion-parameter hybrid state-space language model, developed in-house, engineered specifically to address the complex linguistic landscape of India. The model’s capacity to handle code-mixing, diverse dialects, regional accents, and noisy audio environments positions it as a robust solution for real-world applications.

The significance of Saaras V4 extends beyond its technical prowess, fundamentally reshaping how users and industries interact with digital platforms in India. For the vast non-English speaking population, this model dramatically lowers barriers to digital inclusion, enabling greater access to essential services in e-commerce, fintech, education, and government. By natively understanding and processing speech in local languages, Saaras V4 transforms user experience from a "patched-on" translation to a genuinely intuitive interaction, fostering trust and increasing participation in the digital economy. This is particularly crucial in a country where the Indian voice AI market, valued at approximately USD 153 million in 2024, is projected to surge to USD 957 million by 2030, exhibiting a compound annual growth rate (CAGR) of around 35.7%. The broader speech and voice recognition market in India is also set for exceptional growth, from USD 366.3 million in 2025 to USD 954.3 million by 2030, with a CAGR of 21.1%.

For enterprises, Saaras V4 offers tangible operational efficiencies and enhanced customer engagement. Its ability to provide five distinct output modes—verbatim, normalized transcription, code-mixed text, transliteration, and direct English translation—from a single audio input streamlines workflows for applications like call analytics, voice agents, and customer service. This eliminates the need for multiple post-processing steps, reducing complexity and potential error cascades. Furthermore, the model's keyterm prompting feature, allowing up to 50 domain-specific terms to bias recognition, ensures higher accuracy for industry-specific jargon, brand names, and proper nouns. Its low-latency streaming, with a reported time-to-first-token below 150 milliseconds, makes it suitable for real-time conversational AI, while its batch API can process multi-minute recordings in under a second.

Sarvam AI, founded in August 2023 by Dr. Vivek Raghavan and Dr. Pratyush Kumar, has rapidly positioned itself at the forefront of India's "Sovereign AI" initiative. The company, backed by approximately US$41 million in funding, was selected by India's Ministry of Electronics and Information Technology (MeitY) in April 2025 to develop indigenous foundational models under the IndiaAI Mission. This strategic focus contrasts with global AI giants like Google Cloud Speech, Microsoft Azure Speech, and OpenAI Whisper, which, while capable in global English, often exhibit significant performance degradation when handling the nuances of Indian languages and code-mixing. While competitors like Reverie cover 11+ Indian languages, and Mihup shows strong real-call benchmarking for select languages like Hindi, Saaras V4 distinguishes itself by offering State-of-the-Art (SOTA) performance across *all* 22 Indian languages, coupled with industry-leading accuracy on seven global English datasets. The model's low language identification error rate of 5.22% across all 22 languages (2.9% for the top 10) further highlights its sophisticated understanding of India's linguistic diversity.

Looking ahead, Saaras V4 is poised to accelerate India's digital transformation. Its robust capabilities will likely drive deeper integration of voice AI across various sectors, from automating call centers to enabling more intuitive voice interfaces for rural populations accessing government services or financial tools. Sarvam AI's broader strategy, encompassing solutions like Vision 2.1 for document intelligence, suggests a concerted effort to build a comprehensive, full-stack AI ecosystem tailored for Indian enterprises and public services. This commitment to "Deep Intelligence" over generic AI, focusing on India's unique linguistic and infrastructural realities, will continue to foster technological sovereignty and empower a new generation of developers and startups through programs like the Sarvam Startup Program. The ability of Saaras V4 to make technology truly speak India's diverse languages is not just an incremental improvement; it is a foundational step towards a more inclusive, digitally empowered future for the nation.