AI & ML interests
None defined yet.
Recent Activity
Hmar Heritage Foundation
The Hmar Heritage Foundation is an independent community organization dedicated to building open-access digital infrastructure, linguistic datasets, and cultural archives for the Hmar language (hmr, ISO 639-3). Our mission is to ensure that indigenous and regional languages of Northeast India are preserved, standardized, and represented as first-class citizens in modern computational systems.
Linguistic Datasets & Corpora
We curate, harmonize, and maintain open-access datasets for the Hmar language and the Zo language family on Hugging Face:
| Dataset | Scope / Records | Description | Link |
|---|---|---|---|
sentences |
101,867 train sentences (~2.5M words) | Multi-register sentence corpus and evaluation benchmark for language modeling. | sentences |
paragraphs |
45,103 sequences (~4.37M words) | Long-context multi-register corpus (up to 512 tokens) validated with hmaraniam. |
paragraphs |
unigrams |
52,564 unique tokens | Frequency-weighted vocabulary compiled across 2.73M+ words. | unigrams |
wordlist |
43,509 lexical entries | Structured digital lexicon compiled from 5 lexicographical sources. | wordlist |
zo-bible |
30,974 verse anchors | Sentence-aligned parallel Bible corpus across 8 Zo speech varieties & English. | zo-bible |
numeral-words |
999,999 parallel rows | Sequential spelled-out numerals in Hmar and English (1 to 999,999). |
numeral-words |
hmingtluon |
10.19M synthetic names | Procedural anthroponyms dataset for NER, identity resolution, and clan genealogy. | hmingtluon |
corpus-archive |
Multi-format vault | Scanned books (PDF), structured texts (JSON), and academic linguistic repositories. | corpus-archive |
zo-cognates |
In preparation (Currently empty) | Comparative sound shifts and lexical correspondences across the Zo language family. | — (Currently empty) |
culture-dump |
In preparation (Currently empty) | Unstructured community web publisher HTML dumps, news archives, and cultural assets. | — (Currently empty) |
Language Models & Evaluation Platforms
Community-developed experimental language models and evaluation spaces:
| Resource | Type | Author / Architecture | Parameters | Description | Link |
|---|---|---|---|---|---|
HmarBERT-mini-v2 |
Model | Donal Muolhoi · 8L / 512H | 29.3M | Curriculum pretraining with deterministic whole-word masking and 7k compact vocabulary. | azinamotoe/HmarBERT-mini-v2 |
HmarBERT-mini |
Model | Donal Muolhoi · 4L / 256H | 16.9M | Ultra-lightweight edge/on-device masked language model (24.5k vocab). | azinamotoe/HmarBERT-mini |
HmarBERT (Base POC) |
Model | Donal Muolhoi · 12L / 768H | 110M | Foundational BERT-base proof-of-concept adapted from MizBERT via cognate vocabulary swaps. | azinamotoe/HmarBERT |
| HmarBERT Arena | Space | Hmar Heritage Foundation | — | Interactive comparative explorer and multi-model consensus arbiter across the trio. | HmarBERT Arena |
hmaraniam |
Space / Engine | Hmar Heritage Foundation | — | Fast heuristic language identification library and web demo for Hmar. | Hmaraniam Space |
| Dolma | Space | azinamotoe | — | Conversational and interactive interface powered by Hmar language technology. | Dolma Space |
📌 Explore the full model suite in the HmarBERT Family Collection.
Institutional Links & Standards
- Official Portal: https://hmarheritage.pages.dev
- Language Standards Policy: Standards & Classification
- GitHub Organization: github.com/hmar-heritage-org
hmaraniamPyPI: pypi.org/project/hmaraniam- Primary ISO 639-3:
hmr| Glottolog:hmar1241 - Language Family: Zo Languages
A Note on Nomenclature & Standards
We have minted DOIs for our datasets because we have a very specific philosophy (and a standard) regarding the Kuki-Chin-Naga labels. We want you scholars and researchers to adopt "Zo" over the colonial label. Not because we are under the assumption that we have created something significant. We are aware that our work so far is pretty basic and we will add more datasets but that requires collaboration not within the team but the community and we are working on that. So please try and respect the "labels". You can read more about that here: https://hmarheritage.pages.dev/standards.
Citation & Attribution
If you use datasets, language tools, or resources from the Hmar Heritage Foundation in your research or applications, please cite:
@misc{hmar_heritage_foundation_2026,
author = {{Hmar Heritage Foundation}},
title = {Open Computational Infrastructure, Corpora, and Linguistic Datasets for the Hmar Language},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/hmar-heritage-org},
keywords = {Hmar, hmr, hmar1241, Zo Languages, Low-Resource NLP, Language Preservation},
note = {Language: Hmar (ISO 639-3: hmr, Glottolog: hmar1241). Language family: Zo Languages.}
}