Datasets
3,216 results
Bangla–English Synthetic Parallel Corpus
Bangla–English Synthetic Parallel Corpus
panlex-meanings
Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…
IN22ConvBitextMining
IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…
undl_ru2en_aligned
Dataset Card for "undl ru2en aligned" More Information needed
Parallel_Dataset
Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:
NL2SH-ALFA
Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…
RosettIA Chanka Quechua — Judicial Parallel Data (redistributable subset)
RosettIA Chanka Quechua — Judicial Parallel Data (redistributable subset)
ChEBI-20-MM
ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is…
JMedBench
Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any…
IPA Phonetic Lexicon (7M words)
IPA Phonetic Lexicon (7M words)
Flickr30k
Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…