हिंदी में पढ़ें —JantaScope हिंदी
AI NEWS

IISc’s SraVaani Pushes Voice AI Beyond India’s 22 Mainstream Languages

Researchers at the Indian Institute of Science (IISc) and ARTPARK@IISc have developed SraVaani-1.0, an open-source automatic speech recognition (ASR) model designed to extend voice technology to dozens of Indian languages and dialects, including many that receive little or no support from mainstream speech-AI systems. The research system covers 65 Indian languages and dialects, shifting the focus of India’s voice-AI race from performance in major languages toward broader linguistic inclusion.

IISc’s SraVaani Pushes Voice AI Beyond India’s 22 Mainstream Languages

By Jeet Nirmal

Source: Moneycontrol

IISc’s SraVaani Expands the Boundaries of Indian Voice AI

India’s artificial intelligence ecosystem has made significant progress in building tools that understand languages such as Hindi, Bengali, Tamil, Telugu and Marathi. Yet the country’s linguistic landscape extends far beyond the 22 languages recognised in the Eighth Schedule of the Constitution.

SraVaani-1.0 is attempting to address that wider challenge.

Developed by researchers associated with the SPIRE Lab at IISc and ARTPARK@IISc, with Google support reported for the broader initiative, SraVaani has been designed as a multilingual automatic speech recognition system with an emphasis on languages that historically have had limited digital speech resources.

According to the research paper, the system covers 65 Indian languages and dialects, many of which lack publicly available or competing ASR systems.

That makes the project significant not simply because of how many languages it recognises, but because of which languages it is trying to bring into the voice-AI ecosystem.

Moving Beyond the Biggest Indian Languages

Much of commercial voice technology naturally concentrates on languages with large populations and substantial quantities of digitised speech.

That creates a difficult cycle for smaller languages. Limited datasets make it harder to train accurate AI models, while the absence of reliable technology can reduce incentives to create more digital content and services in those languages.

SraVaani attempts to break part of that cycle by extending speech recognition into India’s linguistic “long tail.”

Researchers report that the model covers languages and dialects for which other evaluated systems may provide limited or no transcription capability. This could be particularly relevant for communities speaking tribal, regional and other low-resource languages.

More Than 31,000 Hours of Speech Used in Training

A major component of SraVaani’s development is the scale and structure of its training process.

The researchers describe a three-stage approach. First, the system underwent self-supervised pretraining using 31,255 hours of unlabelled speech from the VAANI corpus.

The second stage introduced an audio-image alignment process, allowing the speech encoder to learn from relationships between spoken material and corresponding visual information.

Finally, the system was fine-tuned using 31,263 hours of labelled multilingual Indian speech, compiled from 24 public datasets covering the 65 reported languages and dialects.

The multimodal stage is especially notable because low-resource languages often lack the enormous quantities of labelled speech available for globally dominant languages. Using additional contextual information may therefore help models learn stronger representations from relatively scarce resources.

Why SraVaani Matters for India

The significance of inclusive speech recognition extends well beyond virtual assistants.

Voice interfaces could become particularly useful in areas where typing is inconvenient, literacy levels vary or users are more comfortable communicating in their mother tongue.

Better recognition of low-resource Indian languages could eventually support applications including digital public services, agriculture information platforms, education tools, accessibility technologies, healthcare interfaces and customer-service systems.

Consider a citizen trying to interact with a digital service through a language that conventional speech-recognition systems do not understand. Even if that person owns a smartphone and has internet access, the language barrier can still prevent meaningful access to the service.

In that sense, linguistic inclusion is increasingly becoming part of the broader digital-inclusion challenge.

Open-Source Approach Could Expand Development

Another important feature of SraVaani is its open approach.

The research presents SraVaani-1.0 as an open-source model, giving researchers and developers an opportunity to study and adapt the technology rather than relying entirely on closed commercial speech platforms.

This could be especially important for universities, startups and organisations working with smaller language communities.

Instead of developing an ASR model completely from scratch, teams could potentially fine-tune an existing multilingual foundation using specialised datasets for a particular language, dialect or application.

Recent research involving Mizo illustrates both the opportunity and the challenge: researchers found that fine-tuning SraVaani with curated Mizo speech data substantially improved its performance compared with its zero-shot result, showing that broad multilingual coverage does not eliminate the need for language-specific adaptation.

Accuracy Remains a Critical Challenge

Broad language coverage should not automatically be interpreted as equally strong recognition performance across every supported language.

Speech recognition can be affected by accents, dialect variation, background noise, speaking style, recording quality and differences between training data and real-world speech.

SraVaani’s researchers report competitive performance against several multilingual ASR systems across multiple benchmarks, including strong results on a number of low-resource languages. However, the results also underline that inclusive voice AI remains an evolving research problem rather than a solved one.

For applications such as healthcare, financial services or government benefits, recognition errors could have serious consequences. Real-world deployments would therefore need careful testing and safeguards rather than relying only on headline language-coverage numbers.

From Language Count to Real-World Usability

SraVaani also highlights an important change in how India’s voice-AI competition may be evaluated.

The next phase may not simply be about which system performs best in Hindi, Tamil or other widely digitised languages. Increasingly, developers may be judged on whether their systems can handle regional accents, code-switching, dialects and languages for which large training datasets do not exist.

India’s linguistic diversity makes this an unusually demanding technological problem—but also a major opportunity for AI research.

If systems such as SraVaani can eventually combine broad coverage with dependable real-world accuracy, voice technology could reach communities that have largely remained outside the current AI ecosystem.

Balanced Analysis

SraVaani represents an important research direction because it treats linguistic coverage as a core AI problem rather than an afterthought.

Its open-source nature and focus on low-resource languages could encourage researchers and developers to build specialised systems for communities that commercial AI providers may not prioritise.

At the same time, supporting a language technically is only the first step. Practical usefulness depends on recognition accuracy across speakers, dialects, environments and specialised vocabulary. Continued collection of representative datasets, community participation and language-specific fine-tuning are likely to remain necessary.

SraVaani therefore should not be viewed as the final solution to India’s voice-AI divide. Its larger contribution may be demonstrating that the boundaries of Indian speech technology can extend far beyond the country’s most digitally represented languages.


This article is based on reporting published by Moneycontrol.

Related

More stories

Nvidia Is No Longer Just an AI Chip Giant — Its Push Into AI Models Is Getting Much Bigger

Nvidia is expanding beyond the hardware that powered the generative-AI boom and strengthening its position in AI models, particularly through its Nemotron family and a major deal involving AI startup Poolside. The strategy could give Nvidia greater influence across the entire AI technology stack—from computing infrastructure to the models and agents running on top of it.

AI NEWS

Nvidia Is No Longer Just an AI Chip Giant — Its Push Into AI Models Is Getting Much Bigger

Anthropic’s $2 Trillion IPO Dream Could Rewrite Wall Street Records — Here’s Why Investors Are Watching

Artificial intelligence company Anthropic is moving closer to a potential stock-market debut that could become one of the largest IPOs ever. Investors are reportedly discussing a valuation of around $2 trillion or more, highlighting extraordinary expectations surrounding the company behind the Claude AI models.

AI NEWS

Anthropic’s $2 Trillion IPO Dream Could Rewrite Wall Street Records — Here’s Why Investors Are Watching

Meta’s AI Talent War Heats Up Again as OpenAI Veteran Luke Metz Makes the Switch

The battle for the world’s most sought-after artificial intelligence researchers is intensifying again. Meta has hired veteran AI researcher Luke Metz, adding another former OpenAI talent to its expanding Superintelligence Labs operation. The move highlights how competition between leading AI companies is increasingly being fought not only through bigger models and computing infrastructure, but also through the recruitment of a relatively small group of researchers with experience building front

AI NEWS

Meta’s AI Talent War Heats Up Again as OpenAI Veteran Luke Metz Makes the Switch

Nvidia AI Server Prices Could Rise More Than 15% as Memory Chip Costs Surge

Some of Nvidia's biggest customers have reportedly been informed that prices for servers powered by the company's AI processors could increase by more than 15% in many configurations. The increases are expected to affect systems shipped from early 2027, including platforms using Grace Blackwell and next-generation Vera Rubin technology, as soaring memory costs add pressure to the global AI infrastructure boom. Nvidia has not publicly confirmed the reported increases.

AI NEWS

Nvidia AI Server Prices Could Rise More Than 15% as Memory Chip Costs Surge

OpenAI Warns of Growing AI-Powered Cyber Threats as Safety Concerns Intensify

OpenAI is warning that rapidly advancing artificial intelligence could dramatically increase the speed, scale and autonomy of cyberattacks. The company says newer AI systems are becoming increasingly capable at cybersecurity tasks, creating a dual-use challenge: the same technology that can discover vulnerabilities and strengthen defenses could also help malicious actors exploit weaknesses. OpenAI has responded by tightening safeguards, expanding monitoring and giving advanced cyber capabilities

AI NEWS

OpenAI Warns of Growing AI-Powered Cyber Threats as Safety Concerns Intensify

OpenAI Cuts GPT-5.6 Sol Developer Pricing by More Than 20% as AI Competition Intensifies

OpenAI has reduced developer pricing for its flagship GPT-5.6 Sol model by more than 20%, lowering the cost of accessing the company's frontier AI through its API. The promotional pricing is expected to remain available for at least three months, as OpenAI seeks to make advanced AI workloads more economical amid intensifying competition across the global artificial intelligence market.

AI NEWS

OpenAI Cuts GPT-5.6 Sol Developer Pricing by More Than 20% as AI Competition Intensifies