AI & ML interests

None defined yet.

Recent Activity

azinamotoe  updated a Space 8 days ago
hmar-heritage-org/README
azinamotoe  updated a dataset 14 days ago
hmar-heritage-org/unigrams
azinamotoe  updated a dataset 14 days ago
hmar-heritage-org/sentences
View all activity

Organization Card

Hmar Heritage Foundation

The Hmar Heritage Foundation is an independent community organization dedicated to building open-access digital infrastructure, linguistic datasets, and cultural archives for the Hmar language (hmr, ISO 639-3). Our mission is to ensure that indigenous and regional languages of Northeast India are preserved, standardized, and represented as first-class citizens in modern computational systems.


Linguistic Datasets & Corpora

We curate, harmonize, and maintain open-access datasets for the Hmar language and the Zo language family on Hugging Face:

Dataset Scope / Records Description Link
sentences 101,867 train sentences (~2.5M words) Multi-register sentence corpus and evaluation benchmark for language modeling. sentences
paragraphs 45,103 sequences (~4.37M words) Long-context multi-register corpus (up to 512 tokens) validated with hmaraniam. paragraphs
unigrams 52,564 unique tokens Frequency-weighted vocabulary compiled across 2.73M+ words. unigrams
wordlist 43,509 lexical entries Structured digital lexicon compiled from 5 lexicographical sources. wordlist
zo-bible 30,974 verse anchors Sentence-aligned parallel Bible corpus across 8 Zo speech varieties & English. zo-bible
numeral-words 999,999 parallel rows Sequential spelled-out numerals in Hmar and English (1 to 999,999). numeral-words
hmingtluon 10.19M synthetic names Procedural anthroponyms dataset for NER, identity resolution, and clan genealogy. hmingtluon
corpus-archive Multi-format vault Scanned books (PDF), structured texts (JSON), and academic linguistic repositories. corpus-archive
zo-cognates In preparation (Currently empty) Comparative sound shifts and lexical correspondences across the Zo language family. — (Currently empty)
culture-dump In preparation (Currently empty) Unstructured community web publisher HTML dumps, news archives, and cultural assets. — (Currently empty)

Language Models & Evaluation Platforms

Community-developed experimental language models and evaluation spaces:

Resource Type Author / Architecture Parameters Description Link
HmarBERT-mini-v2 Model Donal Muolhoi · 8L / 512H 29.3M Curriculum pretraining with deterministic whole-word masking and 7k compact vocabulary. azinamotoe/HmarBERT-mini-v2
HmarBERT-mini Model Donal Muolhoi · 4L / 256H 16.9M Ultra-lightweight edge/on-device masked language model (24.5k vocab). azinamotoe/HmarBERT-mini
HmarBERT (Base POC) Model Donal Muolhoi · 12L / 768H 110M Foundational BERT-base proof-of-concept adapted from MizBERT via cognate vocabulary swaps. azinamotoe/HmarBERT
HmarBERT Arena Space Hmar Heritage Foundation — Interactive comparative explorer and multi-model consensus arbiter across the trio. HmarBERT Arena
hmaraniam Space / Engine Hmar Heritage Foundation — Fast heuristic language identification library and web demo for Hmar. Hmaraniam Space
Dolma Space azinamotoe — Conversational and interactive interface powered by Hmar language technology. Dolma Space

📌 Explore the full model suite in the HmarBERT Family Collection.


Institutional Links & Standards


A Note on Nomenclature & Standards

We have minted DOIs for our datasets because we have a very specific philosophy (and a standard) regarding the Kuki-Chin-Naga labels. We want you scholars and researchers to adopt "Zo" over the colonial label. Not because we are under the assumption that we have created something significant. We are aware that our work so far is pretty basic and we will add more datasets but that requires collaboration not within the team but the community and we are working on that. So please try and respect the "labels". You can read more about that here: https://hmarheritage.pages.dev/standards.


Citation & Attribution

If you use datasets, language tools, or resources from the Hmar Heritage Foundation in your research or applications, please cite:

@misc{hmar_heritage_foundation_2026,
  author    = {{Hmar Heritage Foundation}},
  title     = {Open Computational Infrastructure, Corpora, and Linguistic Datasets for the Hmar Language},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/hmar-heritage-org},
  keywords  = {Hmar, hmr, hmar1241, Zo Languages, Low-Resource NLP, Language Preservation},
  note      = {Language: Hmar (ISO 639-3: hmr, Glottolog: hmar1241). Language family: Zo Languages.}
}

models 0

None public yet