Skip to content
Advertisement

100K–1M

595 results

MultimodalTabularText

planetarium

Dataset Card for Planetarium🪐 Planetarium🪐 is a dataset and benchmark for assessing LLMs in translating natural language descriptions…

100K–1M·CC-BY·Parquet
MultimodalTabularText

JMedBench

Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any…

100K–1M·JSON
Text

NTREX

Dataset Description NTREX -- News Test References for MT Evaluation from English into a total of 128 target…

100K–1M·CC-BY-SA·Text (raw)
Text

mmarco-contrastive

mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to…

100K–1M·Apache-2.0·Parquet
AudioMultimodalText

wmt-human-all-TTS

WMT Human + TTS Audio WMT human evaluation data (zouharvi/wmt-human-all) extended with TTS-synthesised source audio, covering 49 language…

100K–1M·CC-BY·Parquet
ImageMultimodalText

COCO-35L

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

100K–1M·Parquet
Text

IN22GenBitextMining

IN22GenBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Gen is a n-way parallel general-purpose multi-domain benchmark dataset for…

100K–1M·CC-BY·Parquet
Text

tatoeba-bitext-mining

Tatoeba An MTEB dataset Massive Text Embedding Benchmark 1,000 English-aligned sentence pairs for each language based on the…

100K–1M·CC-BY·JSON
Text

IndicGenBenchFloresBitextMining

IndicGenBenchFloresBitextMining An MTEB dataset Massive Text Embedding Benchmark Flores-IN dataset is an extension of Flores dataset released as…

100K–1M·CC-BY-SA·Parquet
Advertisement