Skip to content
Advertisement

Open

2,922 results

Text

Translate_datasets

Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…

10M–100M·MIT·Arrow
Text

rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…

1M–10M·Apache-2.0·CSV
Text

finetranslations-sample-ar-en

Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…

100M–1B·CC0·Parquet
Text

GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…

10M–100M·Parquet
Text

AfriDocMT

data ├── document │ ├── Health │ │ ├── dev.csv │ │ ├── test.csv │ │ └── train.csv…

10K–100K·CSV
MultimodalTabularText

panlex-meanings

Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…

10M–100M·CC0·CSV
Text

IN22ConvBitextMining

IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…

1K–10K·CC-BY·Parquet
Text

OpusGnome

OpusGnome

10K–100K·Custom / Research-only·Parquet
Text

Parallel_Dataset

Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:

1M–10M·MIT·JSON
Text

NL2SH-ALFA

Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…

10K–100K·MIT·CSV
Advertisement