हिंदी में पढ़ें —JantaScope हिंदी
AI NEWS

IISc’s SraVaani Pushes Voice AI Beyond India’s 22 Mainstream Languages

Researchers at the Indian Institute of Science (IISc) and ARTPARK@IISc have developed SraVaani-1.0, an open-source automatic speech recognition (ASR) model designed to extend voice technology to dozens of Indian languages and dialects, including many that receive little or no support from mainstream speech-AI systems. The research system covers 65 Indian languages and dialects, shifting the focus of India’s voice-AI race from performance in major languages toward broader linguistic inclusion.

IISc’s SraVaani Pushes Voice AI Beyond India’s 22 Mainstream Languages

By Jeet Nirmal

Source: Moneycontrol

IISc’s SraVaani Expands the Boundaries of Indian Voice AI

India’s artificial intelligence ecosystem has made significant progress in building tools that understand languages such as Hindi, Bengali, Tamil, Telugu and Marathi. Yet the country’s linguistic landscape extends far beyond the 22 languages recognised in the Eighth Schedule of the Constitution.

SraVaani-1.0 is attempting to address that wider challenge.

Developed by researchers associated with the SPIRE Lab at IISc and ARTPARK@IISc, with Google support reported for the broader initiative, SraVaani has been designed as a multilingual automatic speech recognition system with an emphasis on languages that historically have had limited digital speech resources.

According to the research paper, the system covers 65 Indian languages and dialects, many of which lack publicly available or competing ASR systems.

That makes the project significant not simply because of how many languages it recognises, but because of which languages it is trying to bring into the voice-AI ecosystem.

Moving Beyond the Biggest Indian Languages

Much of commercial voice technology naturally concentrates on languages with large populations and substantial quantities of digitised speech.

That creates a difficult cycle for smaller languages. Limited datasets make it harder to train accurate AI models, while the absence of reliable technology can reduce incentives to create more digital content and services in those languages.

SraVaani attempts to break part of that cycle by extending speech recognition into India’s linguistic “long tail.”

Researchers report that the model covers languages and dialects for which other evaluated systems may provide limited or no transcription capability. This could be particularly relevant for communities speaking tribal, regional and other low-resource languages.

More Than 31,000 Hours of Speech Used in Training

A major component of SraVaani’s development is the scale and structure of its training process.

The researchers describe a three-stage approach. First, the system underwent self-supervised pretraining using 31,255 hours of unlabelled speech from the VAANI corpus.

The second stage introduced an audio-image alignment process, allowing the speech encoder to learn from relationships between spoken material and corresponding visual information.

Finally, the system was fine-tuned using 31,263 hours of labelled multilingual Indian speech, compiled from 24 public datasets covering the 65 reported languages and dialects.

The multimodal stage is especially notable because low-resource languages often lack the enormous quantities of labelled speech available for globally dominant languages. Using additional contextual information may therefore help models learn stronger representations from relatively scarce resources.

Why SraVaani Matters for India

The significance of inclusive speech recognition extends well beyond virtual assistants.

Voice interfaces could become particularly useful in areas where typing is inconvenient, literacy levels vary or users are more comfortable communicating in their mother tongue.

Better recognition of low-resource Indian languages could eventually support applications including digital public services, agriculture information platforms, education tools, accessibility technologies, healthcare interfaces and customer-service systems.

Consider a citizen trying to interact with a digital service through a language that conventional speech-recognition systems do not understand. Even if that person owns a smartphone and has internet access, the language barrier can still prevent meaningful access to the service.

In that sense, linguistic inclusion is increasingly becoming part of the broader digital-inclusion challenge.

Open-Source Approach Could Expand Development

Another important feature of SraVaani is its open approach.

The research presents SraVaani-1.0 as an open-source model, giving researchers and developers an opportunity to study and adapt the technology rather than relying entirely on closed commercial speech platforms.

This could be especially important for universities, startups and organisations working with smaller language communities.

Instead of developing an ASR model completely from scratch, teams could potentially fine-tune an existing multilingual foundation using specialised datasets for a particular language, dialect or application.

Recent research involving Mizo illustrates both the opportunity and the challenge: researchers found that fine-tuning SraVaani with curated Mizo speech data substantially improved its performance compared with its zero-shot result, showing that broad multilingual coverage does not eliminate the need for language-specific adaptation.

Accuracy Remains a Critical Challenge

Broad language coverage should not automatically be interpreted as equally strong recognition performance across every supported language.

Speech recognition can be affected by accents, dialect variation, background noise, speaking style, recording quality and differences between training data and real-world speech.

SraVaani’s researchers report competitive performance against several multilingual ASR systems across multiple benchmarks, including strong results on a number of low-resource languages. However, the results also underline that inclusive voice AI remains an evolving research problem rather than a solved one.

For applications such as healthcare, financial services or government benefits, recognition errors could have serious consequences. Real-world deployments would therefore need careful testing and safeguards rather than relying only on headline language-coverage numbers.

From Language Count to Real-World Usability

SraVaani also highlights an important change in how India’s voice-AI competition may be evaluated.

The next phase may not simply be about which system performs best in Hindi, Tamil or other widely digitised languages. Increasingly, developers may be judged on whether their systems can handle regional accents, code-switching, dialects and languages for which large training datasets do not exist.

India’s linguistic diversity makes this an unusually demanding technological problem—but also a major opportunity for AI research.

If systems such as SraVaani can eventually combine broad coverage with dependable real-world accuracy, voice technology could reach communities that have largely remained outside the current AI ecosystem.

Balanced Analysis

SraVaani represents an important research direction because it treats linguistic coverage as a core AI problem rather than an afterthought.

Its open-source nature and focus on low-resource languages could encourage researchers and developers to build specialised systems for communities that commercial AI providers may not prioritise.

At the same time, supporting a language technically is only the first step. Practical usefulness depends on recognition accuracy across speakers, dialects, environments and specialised vocabulary. Continued collection of representative datasets, community participation and language-specific fine-tuning are likely to remain necessary.

SraVaani therefore should not be viewed as the final solution to India’s voice-AI divide. Its larger contribution may be demonstrating that the boundaries of Indian speech technology can extend far beyond the country’s most digitally represented languages.


This article is based on reporting published by Moneycontrol.

Related

More stories

Amazon Explores $8 Billion Financing Structure for Nvidia AI Chips

Amazon is reportedly exploring an unusual financing arrangement involving about $8 billion worth of Nvidia's advanced Grace Blackwell AI chips. The proposed structure would transfer thousands of chips into a special-purpose vehicle backed by outside investors, with Amazon continuing to use the processors through a lease arrangement.

AI NEWS

Amazon Explores $8 Billion Financing Structure for Nvidia AI Chips

ChatGPT Adds AI-Powered Virtual Try-On, Letting Shoppers Preview Clothes on Themselves

OpenAI has expanded ChatGPT’s shopping capabilities with an AI-powered virtual try-on feature that generates previews of users wearing clothing and accessories. Users can upload a selfie, try items surfaced in ChatGPT shopping results, or provide their own product image, while a new Favorites feature allows products to be saved for later.

AI NEWS

ChatGPT Adds AI-Powered Virtual Try-On, Letting Shoppers Preview Clothes on Themselves

Google Unveils Gemini 4 Argon, New Frontier AI Model Built for Complex Work and Cyber Defense

Google has introduced Gemini 4 Argon, its new frontier artificial-intelligence model designed for long, complex workflows across software engineering, finance, legal work and cybersecurity. The model is initially being made available to a limited group of trusted cyber defenders rather than the general public, as Google takes a phased approach to deployment and safety testing.

AI NEWS

Google Unveils Gemini 4 Argon, New Frontier AI Model Built for Complex Work and Cyber Defense

Broadcom Could Lend Anthropic Up to $42 Billion as AI Infrastructure Spending Accelerates

Broadcom has agreed to provide Anthropic with access to as much as $42 billion in financing for infrastructure spending, according to details disclosed in Anthropic's IPO documents. The arrangement deepens an already significant relationship between the semiconductor company and the AI developer as demand for computing capacity continues to rise.

AI NEWS

Broadcom Could Lend Anthropic Up to $42 Billion as AI Infrastructure Spending Accelerates

US Lawmaker Presses Major AI Companies Over Possible Chinese Access to Model Weights

U.S. Representative Ro Khanna has asked several leading American artificial intelligence companies to disclose known attempts by China or other hostile actors to gain unauthorized access to their AI model weights, putting cybersecurity around frontier AI systems under renewed congressional scrutiny.

AI NEWS

US Lawmaker Presses Major AI Companies Over Possible Chinese Access to Model Weights

IndiaAI Mission May Be Recalibrated as GPU Supply and Rising Costs Test Compute Expansion

The government is reportedly considering changes to the IndiaAI Mission's compute strategy after delays in GPU availability and rising hardware costs created a gap between committed and currently accessible capacity. The development highlights the challenges India faces as it tries to build affordable AI infrastructure while remaining dependent on global suppliers for advanced processors.

AI NEWS

IndiaAI Mission May Be Recalibrated as GPU Supply and Rising Costs Test Compute Expansion