Skip to content
Advertisement

BiST 💻 Github Repo English 简体中文 Introduction BiST is a large-scale bilingual translation dataset, with “BiST” standing for Bilingual Synthetic Translation dataset. Currently, the dataset contains approximately 60M entries and will continue to expand in the future. BiST consists of two subsets, namely en-zh and zh-en, where the former represents the source language, collected from public data as real-world content; the latter represents the target language… See the full description on the dataset page:

Source: Hugging Face Hub (Mxode/BiST). Metadata imported from the dataset’s Hub tags.

Advertisement