Skip to content
Advertisement

Multimodal

2,368 results

MultimodalTabularText

qm9

Dataset Card for "QM9" QM9 dataset from Ruddigkeit et al., 2012; Ramakrishnan et al., 2014. Original data downloaded…

100K–1M·Parquet
MultimodalTabularText

cc100-documents

cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…

100M–1B·Parquet
MultimodalTabularText

liars-bench

Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.…

10K–100K·Custom / Research-only·Parquet
MultimodalTabularText

codeforces-cots

Dataset Card for CodeForces-CoTs Dataset description CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive…

100K–1M·CC-BY·Parquet
MultimodalTabularText

SylReg

Dataset for SylReg This repository contains the datasets and alignments associated with the paper Speaker-Disentangled Chunk-Wise Regression for…

10M–100M·CC-BY-NC-SA·Parquet
Advertisement