Skip to content
Advertisement

1K–10K

598 results

ImageMultimodalText

HRVQA-2k

A 2k subset of the validation split of the HRVQA dataset ported to HF for ease-of-use in quick…

1K–10K·CC-BY-NC·Parquet
ImageMultimodalText

RSVQA-LR-2k

A 2k subset of the validation split of the RSVQA LR dataset ported to HF for ease-of-use in…

1K–10K·CC-BY·Parquet
ImageMultimodalText

RSVQA-HR-2k

A 2k subset of the validation split of the RSVQA HR dataset ported to HF for ease-of-use in…

1K–10K·CC-BY·Parquet
Audio

dialogue-episodes

Multi-Language Dialogue Episodes Dataset Dataset Description This dataset contains dialogue episodes with multi-language transcripts and separated…

1K–10K·CC-BY-NC·Audio (folder)
Text

NorwegianCourtsBitextMining

NorwegianCourtsBitextMining An MTEB dataset Massive Text Embedding Benchmark Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian…

1K–10K·CC-BY·Parquet
Text

Maori_English_New_Zealand

Dataset for Translation from Maori to English The source of this dataset is scraped from the website TEARA.…

1K–10K·Other·CSV
Audio

StreamUni

The training dataset for the paper 'StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model' Model…

1K–10K·Apache-2.0·Audio (folder)
Advertisement