Skip to content
Advertisement
MultimodalTabularText

English–Zomi Parallel Corpus (1.78M)

English–Zomi Parallel Corpus (1.78M)

English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: Tatoeba v20230412.

Source: Hugging Face Hub (ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412). Metadata imported from the dataset’s Hub tags.

Advertisement