Creating Speech Datasets and Models for underrepresented Indian Languages to foster inclusive speech technology.
Disparities in speech technology support highlight the urgent need for inclusive initiatives.
Most speech technology products support well-resourced languages like English and Hindi, leaving many Indian languages behind due to limited commercial support.
The primary obstacle is the scarcity of sufficient speech datasets, especially for non-scheduled Indo-Aryan and Dravidian languages, and scheduled languages from Tibeto-Burman and Austro-Asiatic language families.
Initiatives are needed to focus on collecting and curating speech data for underrepresented languages to ensure equitable access to speech technology.
Creating transcribed speech datasets for Indian languages.
Developing robust speech models for these languages.
Focusing on under-resourced languages from Tibeto-Burman and Austro-Asiatic families.
Gathering 100 hours of speech data from at least 10 underrepresented languages within each major language family.
Creating comprehensive phone sets for each language under study to facilitate accurate phonetic representation.
Building baseline Automatic Speech Recognition (ASR) systems for each language to enable practical applications.
Constructing language models to improve the accuracy and fluency of speech recognition systems.
To build pre-trained models based on the data collected in the project for each language family under study to enable further fine-tuning.
Making datasets and pre-trained models publicly available through appropriate platforms.
Using CC BY-SA-NC 4.0 license for the dataset to encourage sharing and adaptation for non-commercial purposes.
Employing AGPL v3 license for the models to ensure open-source access and community-driven improvements.
Focusing on the four major language families of India to ensure broad coverage.
Including at least 10 underrepresented languages in each family for data collection.
Aiming to collect approximately 1000 hours of transcribed speech per language.
The SpeeD-IL project is poised to make a significant impact on the landscape of speech technology in India, driving inclusivity, fostering innovation, and empowering communities through accessible language technologies.
© 2022-26 UnReaL-TecE LLP. All rights reserved.
SpeeD-IL: Bridging the Speech Technology Gap in India