IISc Launches SraVaani Voice AI Model for 65 Indian Languages
| General Studies Paper III: Artificial Intelligence, Scientific Innovations & Discoveries |
Why in News?
Recently, the Indian Institute of Science (IISc) in Bengaluru launched SraVaani, an open-source speech-recognition model.

What is Voice AI Model “SraVaani”?
- About: SraVaani (SraVaani-1.0) is a multilingual Automatic Speech Recognition (ASR) model designed for Indian languages and dialects.
- ASR is an AI technology that transforms human speech into written text.
- Launched By: SraVaani-1.0 was released as an open-source AI model by the ARTPARK–IISc ecosystem.
- The model is publicly hosted under ARTPARK-IISc.
- Its Hugging Face model is released under an MIT licence.
- Developed By: The model was developed by the SPIRE Lab in collaboration with ARTPARK and with support from Google.
- Need: India has 700+ languages and thousands of dialects, but speech-AI systems traditionally support only a limited linguistic subset.
- SraVaani addresses this digital language divide by developing recognition capability for languages.
- Features:
- Language Coverage: The research describes SraVaani-1.0 as covering 65 Indian languages and dialects.
- These include 20 scheduled languages as well as 45 regional, low-resource and tribal languages.
- Its supports 10 scripts and includes automatic language identification.
- Its coverage includes major languages such as Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada and Malayalam, alongside regional and low-resource languages such as Gondi, Bhili, Garhwali, Marwari, Tulu, Kokborok and Wancho.
- Massive Speech Dataset: The model was pretrained using 31,255 hours of unlabelled speech from the VAANI corpus.
- Its final supervised training used 31,263 hours of labelled multilingual Indian speech compiled from 24 public datasets.
- The VAANI programme itself spans a much broader linguistic landscape, with ARTPARK reporting 31.2k hours, 109 languages, 165 districts and 156,000+ speakers in its wider data initiative.
- Core Technology: SraVaani uses the FastConformer architecture for speech processing.
- Its training follows three stages: self-supervised speech pretraining, audio-image representation alignment, and supervised multilingual fine-tuning.
- The final system uses a Hybrid Token-and-Duration Transducer (TDT)-CTC decoder.
- Multimodal Learning Feature: A distinctive feature is its audio-image alignment stage.
- The system uses paired speech and images to develop richer representations of spoken content.
- This is particularly relevant for low-resource languages, where conventional labelled speech data are limited.
- The VAANI corpus provides approximately 11 million audio-image pairs for this purpose.
- Performance: SraVaani was evaluated against three state-of-the-art multilingual ASR systems across eight benchmarks.
- These include datasets such as Common Voice, FLEURS, IndicTTS, Kathbath, RESPIN, GramVaani, MUCS and VAANI.
- The research reports the lowest Word Error Rate (WER) across a large number of language-dataset combinations.
- Language Coverage: The research describes SraVaani-1.0 as covering 65 Indian languages and dialects.
- Significance: SraVaani addresses India’s digital linguistic divide by bringing speech-AI capability to languages traditionally.
- It can potentially support digital governance, education, accessibility, voice interfaces, documentation and language preservation.
Other Indigenous Multilingual Speech-AI Models
- IndicConformer: IndicConformer is an AI4Bharat speech-recognition model based on the Conformer architecture.
- Its model has about 30 million parameters and is designed for real-time ASR across Indian languages.
- AI4Bharat’s broader objective is to build ASR systems covering all 22 constitutionally recognised languages.
- IndicWav2Vec: IndicWav2Vec is another AI4Bharat contribution to multilingual Indian speech recognition. It uses self-supervised learning, reducing dependence on manually transcribed speech.
- Sarvam AI (Saaras & Bulbul): Saaras provides streaming speech-to-text for 22 scheduled languages with code-mixing support, while Bulbul delivers natural text-to-speech across multiple Indian languages.
- IndicVoices: It is a major speech-data initiative funded by Bhashini, MeitY, with support from EkStep Foundation and Nilekani Philanthropies.
- By October 2025, it had collected 7,348 hours of speech from 16,237 speakers, covering 22 languages, and 145 districts.
- Kathbath: It is an important dataset in India’s speech-AI research ecosystem rather than a standalone commercial ASR model.
- Shrutilipi: It is another speech-data resource used within AI4Bharat’s ASR research. Its importance lies in providing additional Indian-language audio and transcription resources.
Government Initiatives
- National Language Translation Mission: The National Language Translation Mission (NLTM) is the policy foundation behind India’s multilingual language-technology push.
- It seeks to reduce language barriers through AI-based translation, speech recognition, text-to-speech, NLP, OCR and transliteration.
- Digital India BHASHINI: BHASHINI (BHASa INterface for India) was launched in 2022 as an AI-led language-technology platform.
- It aims to make digital content and public services accessible in Indian languages, including through voice-based interfaces.
- It is implemented by the Digital India BHASHINI Division under MeitY.
- BHASHINI Language Services: BHASHINI integrates multiple technologies rather than relying on one AI model.
- Its ecosystem includes speech-to-text, text-to-speech, machine translation, transliteration, OCR and NLP.
- BhashaDaan: BhashaDaan addresses the crucial problem of insufficient Indian-language training data.
- Citizens can contribute voice recordings, transcriptions, translations and labelled information through initiatives such as Bolo India, Suno India, Likho India and Dekho India.
- BHASHINI Samudaye: BHASHINI Samudaye was developed to bring together language experts, universities, civil society, data practitioners and technology stakeholders.
- By March 2026, the programme had onboarded more than 10,000 contributors for dataset, annotation and ecosystem activities.
- BHASHINI Sanchalan: BHASHINI Sanchalan focuses on integrating multilingual AI into government systems and public-service delivery.
- It supports voice-first interfaces, translation, domain-specific language models and terminology standardisation.
- National Hub for Language Technologies: The National Hub for Language Technologies (NHLT) operates within the BHASHINI ecosystem as a large-scale language-AI infrastructure layer.
- In March 2026, the government reported that it operated with 350+ optimised AI models and used a vendor- and cloud-agnostic AI Sovereign Cloud.
- Sectoral Adoption: Multilingual AI is being expanded into specific sectors. Nyaya Setu uses BHASHINI’s ASR, NLP and conversational-AI capabilities for voice-first legal assistance.
- GeM is also collaborating with BHASHINI for multilingual public procurement, including voice technologies and domain-specific language models.
Frequently Asked Questions (FAQs):
1. What is the IISc SraVaani Voice AI Model?
SraVaani-1.0 is an open-source multilingual automatic speech-recognition model for 65 Indian languages and dialects.
2. Who developed the SraVaani Voice AI Model?
It was developed by researchers at IISc’s SPIRE Lab, with ARTPARK and associated collaborators, including Google-supported language-data work.
3. Why is SraVaani important for India’s Voice AI ecosystem?
It expands speech-AI coverage to low-resource and tribal languages, addressing India’s major linguistic-data and digital-inclusion gap.
4. Which Indian languages does the SraVaani model support?
SraVaani supports 65 Indian languages and dialects, including major, regional, low-resource and tribal languages.
5. How does the SraVaani Voice AI Model work?
It uses FastConformer, self-supervised speech learning, audio-image alignment, and supervised multilingual training with a TDT-CTC decoder.
Disclaimer: Information in this article is based on official announcements and public records. Regulations and implementation details may evolve over time.