Government jobs • Exam updates • PreparationIndependent information portal
Current Affairs

IISc Bengaluru Releases 'SraVaani', a Multilingual Indian Speech-Recognition AI Model

On 13 August 2026, the Indian Institute of Science (IISc), Bengaluru unveiled 'SraVaani', an open-source multilingual speech recognition model developed by SPIRE Lab in partnership with ARTPARK and supported by Google. SraVaani can transcribe spoken words into text across 65 Indian languages and dialects, including 20 scheduled languages and 45 regional ones like Garo, Angika, and Tulu, supporting ten scripts. It features automatic language identification, eliminating manual language selection during recognition. The model is publicly available under the MIT license on the Hugging Face platform, promoting broad usage and modification. This innovation leverages over 31,000 hours of speech data collected from 156,000 people across 165 districts under Project Vaani, enabling advanced speech technology for resource-constrained Indian languages.

On this page

Key Facts

  • Release Date: 13 August 2026
  • Developed by: SPIRE Lab at IISc Bengaluru, in collaboration with ARTPARK, supported by Google
  • Model Name: SraVaani
  • Languages and Dialects Covered: 65 total — comprising 20 Scheduled Indian languages plus 45 regional languages and dialects including Garo, Angika, Chakma, Kokborok, Tulu, Bundeli, and Bajjika
  • Scripts Supported: 10 scripts
  • Features: Automatic language identification to eliminate the need for manual language selection before speech recognition
  • Availability: Open-source under the MIT license, available on the Hugging Face platform
  • Training Data Source: Project Vaani speech corpus containing over 31,000 hours of natural conversational recordings from 156,000 individuals across 165 districts in 28 Indian states
  • Performance: Achieved a word error rate of 9.5% on the Garo language, a significant improvement over the next-best system with a 69.4% error rate

Background & Context

Speech recognition technology converts spoken language into text, serving as a foundational technology for AI applications including natural language processing and assistive technologies. India’s rich linguistic diversity, with many resource-constrained languages lacking large digital datasets, poses a unique challenge. Project Vaani is a pioneering data collection initiative deploying a geo-centric strategy, gathering speech data representative of the country’s extensive dialectal and regional variations. The collaboration led to the creation of SraVaani, a state-of-the-art multilingual automatic speech recognition (ASR) model that integrates extensive, real-world speech data to enhance recognition accuracy across diverse languages and dialects.

Why This Matters for Exams / Exam Relevance

The release of SraVaani is a critical development in India’s AI and technological landscape, relevant to competitive examinations focusing on current affairs, technology, and linguistic diversity. Important points include the role of IISc Bengaluru and government initiatives, the scope and scale of language coverage, integration of Project Vaani's district-wise speech data, and the significance of open-source licensing fostering innovation. Understanding these aspects reflects wider themes of digital inclusion, technological empowerment for resource-limited languages, and India's leadership in pioneering multilingual AI models.

Points to Remember

  • SraVaani was launched on 13 August 2026 by IISc Bengaluru's SPIRE Lab along with ARTPARK and Google support.
  • The model covers 65 Indian languages and dialects, spanning 20 scheduled and 45 regional languages.
  • It supports ten scripts and includes automatic language identification to simplify usage.
  • The model is trained on over 31,000 hours of speech data collected via Project Vaani, emphasizing real conversational speech and regional diversity.
  • SraVaani is open-source under the permissive MIT license, available through the Hugging Face AI platform.
  • It achieves notably low word error rates, such as 9.5% for the Garo language, demonstrating superior performance on resource-scarce languages.
  • Project Vaani’s geo-centric collection strategy ensures rich dialectal and sociolinguistic representation beyond just named languages.
  • The initiative supports broader AI research and assistive technologies for multilingual and underrepresented Indian languages.
← Back to Current Affairs