Generating inflectional wordforms and morphological features for 10+ Indian languages, facilitating synthetic dataset creation and system development.
Generate all inflectional wordforms from a given lemma or wordform.
Produce complete morphological features for each generated form.
Optionally output the specific rule applied during form generation.
Built on regular expressions, the custom rule-expression language simplifies rule editing.
Currently supports more than 10 Indian languages, with ongoing expansion.
Create large-scale synthetic datasets for training NLP models.
Develop generative systems adaptable to various hardware configurations.
A widely spoken language in North India.
Dominant in West Bengal and Bangladesh.
Spoken in the Assam region.
Predominant in Maharashtra.
Spoken in Odisha.
Official language of Nepal and spoken in Sikkim.
An ancient and classical language.
Spoken in the Punjab region.
The MorphGen language is a custom language designed on top of regular expressions for describing morphological rules in Indian languages. It provides a concise and intuitive way to specify complex patterns and transformations, leveraging the power of regular expressions while adding higher-level abstractions tailored for morphological analysis and suited for non-technical experts.
High precision in morphological generation.
Easy to adapt to new languages and rules.
Fast generation of morphological forms.
Rule-based System for synthetic data generation
© 2022-26 UnReaL-TecE LLP. All rights reserved.
MorphGen: Morphological Generators for Indian Languages