Skip to content
Advertisement

CSV

173 results

Text

ArzEn-MultiGenre

ArzEn-MultiGenre: A Comprehensive Parallel Dataset Overview ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection…

10K–100K·CC-BY·CSV
Text

rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…

1M–10M·Apache-2.0·CSV
Text

AfriDocMT

data ├── document │ ├── Health │ │ ├── dev.csv │ │ ├── test.csv │ │ └── train.csv…

10K–100K·CSV
MultimodalTabularText

panlex-meanings

Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…

10M–100M·CC0·CSV
Text

NL2SH-ALFA

Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…

10K–100K·MIT·CSV
ImageMultimodalTabular

ChEBI-20-MM

ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is…

10K–100K·MIT·CSV
Text

tanglish-tamil

Translation of Tanglish to tamil Source: karky.in To use python import datasets s = datasets.load dataset('Deepakvictor/tanglish-tamil') print(s)…

<1K·OpenRAIL·CSV
MultimodalTabularText

panlex-meanings

Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…

100M–1B·CC0·CSV
Text

PHINC

Abstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities,…

10K–100K·CC-BY·CSV
Advertisement